الباحثون

Zheng Zhang

المنشورات 13

نسخة أولية وصول مفتوح

AdSpark: A Large-Scale Dataset and Benchmark for Product-Centric Advertisement Video Generation

Zhifei Yang, Zhao Jiang, Keyang Lu وآخرون · 2026

Product-centric advertisement video generation aims to create promotional videos that preserve fine-grained product identity while presenting selling points through coherent multi-shot narratives. However, this emerging task remains underexplored due to the lack of large-scale advertisement-specific datasets and compre …

نسخة أولية وصول مفتوح

ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents

Yupeng Su, Jiayi Tian, Zheng Zhang وآخرون · 2026

Long-horizon LLM agents act on an append-only interaction history that is re-sent to the model at every step, so the context and its cost grow with steps until the sessions exceed the context window. Existing methods manage the context through context requirement prediction, relying on additional model calls, heuristic …

نسخة أولية وصول مفتوح

ALoDLM: Adaptively Looped Diffusion Language Models

Liancheng Fang, Zhuowei Li, Youngeun Kim وآخرون · 2026

Diffusion language models (DLMs) enable fast generation by predicting multiple tokens in parallel, but their practical adoption remains limited by a persistent quality gap relative to comparably sized autoregressive (AR) models. We attribute this gap to a computation-difficulty mismatch: within a partially observed seq …

نسخة أولية وصول مفتوح

Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation

Rui Li, Liyang He, Zheng Zhang وآخرون · 2026

On-policy self-distillation (OPSD) supplies dense token-level feedback on trajectories sampled from the student's own policy, a richer training signal than the outcome-level rewards of reinforcement learning. This feedback comes from a teacher conditioned on a full reference solution unavailable to the student. The ref …

نسخة أولية وصول مفتوح

Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

Yu Li, Zheng Zhang, Xin Liu وآخرون · 2026

Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more …

نسخة أولية وصول مفتوح

Does the VGGT Family Need All Its Layers?

Fengyi Zhang, Holger Caesar, Xiangyu Sun وآخرون · 2026

Which layers of a feed-forward geometry model are needed to preserve both camera poses and dense 3D structure? We study layer redundancy in VGGT, $π^3$, and VGGT-$Ω$: 3,018 pruned configurations, scored on seven camera-pose and dense-geometry metrics across four indoor and outdoor datasets. Four findings follow: (i) Re …

نسخة أولية وصول مفتوح

Train Ahead, Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models

Zheng Zhang, Xinyue Tan, Lufei Li وآخرون · 2026

On-policy self-distillation (OPSD) improves large language models by letting a self-teacher with privileged information provide dense token-level supervision on the model's own trajectories. Yet existing methods typically construct the self-teacher from the current, initial, or slowly averaged policy state, leaving the …

نسخة أولية وصول مفتوح

From Judgment Quality to Downstream Utility: Rethinking LLM-as-a-Judge for Open-Ended Tasks

Zheng Zhang, Lufei Li, Xinyue Tan وآخرون · 2026

LLM-as-a-Judge is increasingly used to evaluate policy responses on open-ended tasks that lack ground-truth answers. Existing work often directly converts the resulting judgments into reward signals for policy training, paying limited attention to intrinsic judgment quality and largely restricting the use of Judges to …

نسخة أولية وصول مفتوح

From Learner Behavior to Reusable Skills for Effective and Efficient Learner Simulation

Zijian Chen, Zheng Zhang, Miao Jia وآخرون · 2026

Learner simulation aims to reproduce how a particular learner behaves on new tasks. Although Large Language Models (LLMs) can generate increasingly fine-grained learning behaviors, existing approaches often need to repeatedly process a growing interaction history to reconstruct the learner. This introduces additional c …

نسخة أولية وصول مفتوح

Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

Xi Xiao, Tianchen Zhao, Youngeun Kim وآخرون · 2026

Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent toke …

نسخة أولية وصول مفتوح

Bench2Dex: Benchmarking Visuo-Tactile Bimanual Dexterous Manipulation Across Dexterous Hands

Zhenjie Yang, Yideng Zhang, Dongjie Zhang وآخرون · 2026

Tactile sensing provides contact information that can be difficult to infer from vision alone, but tactile hardware for dexterous hands has not converged to a common design. Dexterous hands differ in finger structure, contact surfaces, and sensor layouts, while simulated tactile signals still differ from measurements p …

المؤلفون المشاركون