الباحثون

Tong Zheng

المنشورات 5

نسخة أولية وصول مفتوح

From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery

Xinglin Wang, Zishen Liu, Tong Zheng وآخرون · 2026

Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontie …

نسخة أولية وصول مفتوح

CARE: Certifying Acceleration for Vision-Language-Action Inference

Rui Liu, Tong Zheng, Jindong Gu وآخرون · 2026

While vision-language-action (VLA) models have advanced rapidly, running them at every control step remains expensive. Prior work accelerates VLA inference using techniques like action chunking and visual-token pruning, typically evaluating based on latency and average task success. However, acceleration may discard in …

نسخة أولية وصول مفتوح

Can Terminal Agents Trust Their Own Verification? Diagnosing and Improving Self-Verification

Yingfeng Luo, Shaowei Wei, Daixin Wang وآخرون · 2026

Terminal agents rely on self-verification to assess and correct their solutions as they solve tasks through interaction with command-line environments. Yet how trustworthy such self-verification is remains poorly understood. To investigate this question systematically, we introduce a diagnostic framework that identifie …

نسخة أولية وصول مفتوح

Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL

Tong Zheng, Skylar Zhai, Zhan Cheng وآخرون · 2026

Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this …

نسخة أولية وصول مفتوح

Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL

Youling Huang, Tiankuo Xu, Jiaji Liu وآخرون · 2026

Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-stage policies often fail to solve sampled tasks, leaving little useful reward signal for learning. To mitigate this problem, we use on-policy distillation (OPD) to provid …

المؤلفون المشاركون