الباحثون

Yu Li

المنشورات 14

نسخة أولية وصول مفتوح

Recurrent Self-Improvement: Dynamic Cross-Loop On-Policy Distillation for Looped Language Models

Yi Wang, Rui Qian, Yu Li وآخرون · 2026

Looped Language Models (LoopLMs) offer a parameter efficient approach to scaling reasoning by reusing shared parameters across recurrent computation steps. Despite their promise, effective post-training of LoopLMs remains challenging. Existing approaches either provide reward based supervision that is sparse or costly …

نسخة أولية وصول مفتوح

Towards Efficient Robotic Manipulation Models with Self-Recursive Pruning

Zijia Chen, Yuenan Hou, Yu Li وآخرون · 2026

Network pruning can reduce parameter redundancy in robotic policies. However, generic pruning criteria are tailored for image recognition tasks and commonly designed to preserve weight magnitude, local reconstruction, or language-model likelihood rather than closed-loop action behavior. Directly applying these pruning …

نسخة أولية وصول مفتوح

Beyond Perturbation Magnitude: Direction-Dependent Responses in Multimodal Geometric Representations

Yongsheng Luo, Wengan He, Yu Li وآخرون · 2026

Geometric alignment scores based on Gram determinants provide a compact way to model higher-order consistency among modalities, yet how such scores respond to modality degradation is poorly understood. This paper asks whether the response of a multimodal geometric score is determined primarily by the magnitude of the p …

نسخة أولية وصول مفتوح

CIPO: Counterfactual Imagination Policy Optimization for Adaptive Tool Granularity Selection

Yu Li, Yunlu Wan, Zijian Zhu وآخرون · 2026

Large language model (LLM) agents solve complex tasks through multi-step interactions with external tools. These interactions often contain recurring local tool sequences. Treating such sequences as composite "Skills" can shorten tool-use trajectories and reduce repeated low-level decisions. However, when atomic tools …

نسخة أولية وصول مفتوح

Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps

Jiaxin Zhang, Xiangyu Peng, Qinglin Chen وآخرون · 2026

Reinforcement learning for long-horizon agents relies on purely retrospective training signals: credit is assigned only after observing environmental consequences, leaving the agent's belief at action time invisible to the gradient. We introduce Prospective Hindsight (PH), a self-calibrating training principle that aug …

نسخة أولية وصول مفتوح

Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents

Yu Li, Guangfeng Cai, Long-Fei Li وآخرون · 2026

Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not ex …

نسخة أولية وصول مفتوح

Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

Yu Li, Zheng Zhang, Xin Liu وآخرون · 2026

Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more …

نسخة أولية وصول مفتوح

Learning from synthetic photorealistic raindrop for single image raindrop removal

Zhixiang Hao, Shaodi You, Yu Li وآخرون · 2026 · 10.1109/iccvw.2019.00534

Raindrops adhered to camera lens or windshield are inevitable in rainy scenes and can become an issue for many computer vision systems such as autonomous driving. Because raindrop appearance is affected by too many parameters, therefore it is unlikely to find an effective model based solution. Learning based methods ar …

نسخة أولية وصول مفتوح

RoboFFT: Finetuning generative robot policy via online reinforcement learning with forward process

Yu Li, Shenghe Hu, Yuhan Wang وآخرون · 2026

Generative models, such as diffusion and flow-based models, have shown strong promise for robot policy learning by capturing complex and multimodal action distributions from demonstrations. However, policies trained solely with imitation learning often suffer from imperfect demonstrations and distributional shifts, whi …

نسخة أولية وصول مفتوح

CT-OPD: Counterfactual Trace On-Policy Distillation for Diffusion Vision-Language Models

Long Qian, Bingke Zhu, Jiaqi Wei وآخرون · 2026

Diffusion vision-language models generate answers by gradually resolving masked tokens, making accurate conditional prediction in partially resolved states central to post-training. Masking completed answers yields coherent contexts and targets, but prescribed masks do not reflect the model's reveal decisions. Its traj …

نسخة أولية وصول مفتوح

VideoGen-Agent: Reinforcing Video Generation Agents

Binxu Li, Haoyi Duan, Yuhui Zhang وآخرون · 2026

Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, specific identities, physical consistency, or ordered events. In this paper, we present VideoGen-Agent, a multimodal agent …

المؤلفون المشاركون