الباحثون

Lei Liu

المنشورات 6

نسخة أولية وصول مفتوح

From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL

Yitong Qiao, Tiantian He, Lei Liu وآخرون · 2026

Reinforcement learning with a verifiable reward (RLVR) offers a scalable approach to training language-model agents, yet sparse outcome rewards can leave early training with little signal for policy improvement. We identify an On-Policy Acceleration Phenomenon: in our main comparisons, RLVR initialized with on-policy d …

نسخة أولية وصول مفتوح

OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

Dingyuan Dai, Heli Qi, Lei Liu وآخرون · 2026

Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and producing verifiable results. We thus introduce OSWorld-Science, a benchmark and evaluation en …

نسخة أولية وصول مفتوح

Can LLMs Value the Right Evidence? Evidence-Value Misalignment in Dynamic Medical Diagnosis

Kehua Feng, Yunsheng Lu, Yitong Qiao وآخرون · 2026

A correct diagnosis reached from insufficient or misleading evidence can pose a clinical hazard, yet outcome-based accuracy may reward such lucky guesses. We call this mismatch between diagnostic decisions and the value of available evidence Evidence-Value Misalignment (EVM). To disentangle evidential grounding indepen …

نسخة أولية وصول مفتوح

MERID: Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis

Lei Liu, Zhaokang Liang, Qingcheng Zeng وآخرون · 2026

Major depressive disorder (MDD) severely impacts daily activities and quality of life. Detecting MDD involves multimodal data, such as interview recordings and sensor measurements. This is particularly challenging, as these heterogeneous modalities often demand distinct, customized prediction pipelines. Existing effort …

نسخة أولية وصول مفتوح

AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research

Kailin Jiang, Lei Liu, Jian Xi وآخرون · 2026

Document rerankers determine what evidence reaches the downstream model in RAG and deep research, yet mainstream rerankers select by relevance matching, and individually relevant documents rarely constitute the complete, complementary, non-redundant set a complex information need demands. Prior work rewards a set by it …

المؤلفون المشاركون