الباحثون

Zhen Xu

المنشورات 5

نسخة أولية وصول مفتوح

Beyond Score Accuracy: Examining the Diagnostic Quality of LLM-Generated Structured Assessment in Higher Education

As Large Language Models (LLMs) are increasingly adopted for automated grading and feedback in higher education, their structured outputs, including multi-dimensional rubric scores, detailed feedback comments, and improvement suggestions, create an appearance of thorough analytic evaluation. This study examines whether …

نسخة أولية وصول مفتوح

AgentPersonaBench: Benchmarking Persona-Driven User Simulation

Jintao Huang, Yifan Wang, Hongyu Shen وآخرون · 2026

We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic b …

نسخة أولية وصول مفتوح

SEIS: Self-Evolving Inference Systems

Zhen Xu, Jingyu Liu, Zongze Li وآخرون · 2026

Inference systems determine how fast and how cheaply language models can be served, so making them faster has direct practical value. However, prior work focuses mostly on optimizing certain parts such as kernels or memory within the large system. In this work, we take a holistic approach and apply agentic self-evoluti …

نسخة أولية وصول مفتوح

LEAP: Learning Efficient Action Proposals For LLM Agents

Zhen Xu, Qizheng Zhang, Gerry Wan وآخرون · 2026

LLM agents are known to be slow in rollouts. An agent completes a task one step at a time. At each step, it reasons and then chooses an action to execute. The next step and action cannot start until the previous one has finished. Speculative decoding accelerates the rollouts at the reason phase by drafting and verifyin …

نسخة أولية وصول مفتوح

On the Divergence of Accuracy and Mechanism Consistency in Time Series World Models

Haochen Zhang, Jiaheng Guo, Zhen Xu وآخرون · 2026

A time series world model (TSWM) predicts a controlled system's state from its observed history and planned actions and exogenous inputs. Current approaches build forecasters with actions as covariates, trained and evaluated on prediction error under the executed plan. Yet world models compare unexecuted plans, but the …

المؤلفون المشاركون