الباحثون

Chongyang Tao

المنشورات 2

نسخة أولية وصول مفتوح

Beyond Entropy: Self-Diagnostic Multi-Role Token Optimization for Video Reasoning

Yudong Han, Yong Wang, Zaiquan Yang وآخرون · 2026

Reinforcement learning with verifiable rewards has substantially advanced multimodal reasoning, yet it remains fundamentally limited by ambiguous token-level credit assignment. While high-entropy token heuristics encourage possibility exploration, naively extending them to video reasoning tends to induce lengthy reason …

نسخة أولية وصول مفتوح

DeFA: Dependency-Guided Failure Attribution for LLM Agents

Bo Deng, Xinlei Zheng, Yi Wei وآخرون · 2026

Errors in LLM agent executions and their visible consequences can be separated by many steps, making decisive-error localization a matter of understanding both step content and step dependencies. We introduce DeFA, a dependency-guided framework for agent failure attribution. DeFA first combines protocol relations and s …

المؤلفون المشاركون