الباحثون

Lei Feng

المنشورات 8

نسخة أولية وصول مفتوح

Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions

Bojun Yang, Haochen Zhou, Zhifang Zhang وآخرون · 2026

Large vision-language models (LVLMs) are increasingly deployed in safety-critical applications, yet they remain vulnerable to backdoor attacks. Defending against such attacks remains costly, as existing methods require either extensive retraining on clean data or per-query intervention at inference time. To address thi …

نسخة أولية وصول مفتوح

CIPO: Counterfactual Imagination Policy Optimization for Adaptive Tool Granularity Selection

Yu Li, Yunlu Wan, Zijian Zhu وآخرون · 2026

Large language model (LLM) agents solve complex tasks through multi-step interactions with external tools. These interactions often contain recurring local tool sequences. Treating such sequences as composite "Skills" can shorten tool-use trajectories and reduce repeated low-level decisions. However, when atomic tools …

نسخة أولية وصول مفتوح

Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents

Yu Li, Guangfeng Cai, Long-Fei Li وآخرون · 2026

Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not ex …

نسخة أولية وصول مفتوح

Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

Yu Li, Zheng Zhang, Xin Liu وآخرون · 2026

Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more …

نسخة أولية وصول مفتوح

Decoupling Token Roles in Autoregressive Pretraining

Suqin Yuan, Runqi Lin, Kevin Qinghong Lin وآخرون · 2026

Autoregressive pretraining increasingly draws on heterogeneous data, making it important to understand how a model learns from an individual token. The next-token prediction objective naturally identifies a token's contribution with its own loss. However, each token is not only a prediction target but also context for …

نسخة أولية وصول مفتوح

Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment

Suqin Yuan, Runqi Lin, Muyang Li وآخرون · 2026

Human-feedback alignment has made language models useful assistants and is commonly described as aligning them with humans. However, the responses people prefer from an AI need not be the responses they themselves would give. We distinguish alignment with human preferences from alignment with human behavior, and show t …

المؤلفون المشاركون