الباحثون

Mengnan Du

المنشورات 7

نسخة أولية وصول مفتوح

From a Prompt to Repertoires: Evolving Functional REpertoires Enable LLM Continual Learning

Fengyuan Liu, Yue Wang, Hangxi Guo وآخرون · 2026

Continual learning remains challenging for large language models, which must enable models to acquire new skills and knowledge without degrading existing capabilities. Existing approaches typically address this challenge by carefully designing how model parameters are updated. In contrast, prompt optimization avoids co …

نسخة أولية وصول مفتوح

Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation

Hangxi Guo, Fengyuan Liu, Yue Wang وآخرون · 2026

Reinforcement learning is commonly used to train language agents in interactive environments, but cannot be directly applied when rewards are unavailable. Recent methods use environmental feedback as privileged context for hindsight self-distillation, but our analysis suggests that simply conditioning the teacher on fe …

نسخة أولية وصول مفتوح

HeadEdit: Calibrating Language Model Behavior Through the Frozen Unembedding Matrix

Zirui He, Haiyan Zhao, Jingyu Hu وآخرون · 2026

Alignment does not eliminate behavioral errors in language models. Models may still refuse benign requests, call unnecessary tools, or yield to false user claims. Current methods mitigate such errors as a computation problem, and rarely explore if the desired behavior is already encoded in the model's representation. M …

نسخة أولية وصول مفتوح

MM-FinEval: A Multi-Task Multimodal Benchmark for Real-World Financial Forecasting

Dong Shu, Yanguang Liu, Huopu Zhang وآخرون · 2026

Financial forecasting from earnings conference calls requires models to reason over complex corporate disclosures, market expectations, and subtle communication signals. However, existing financial benchmarks are often limited to unimodal inputs or single-task settings, making it difficult to evaluate whether multimoda …

نسخة أولية وصول مفتوح

Look Before You Select: Rethinking Vocabulary Sparsification in On-Policy Distillation

Yongliang Miao, Shuang Liu, Yanguang Liu وآخرون · 2026

On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes memory-intensive for long sequences. Existing memory-saving app …

نسخة أولية وصول مفتوح

Faithful Activation Verbalization: Reducing Hallucinations in LLM Representation Interpretation

Haiyan Zhao, Zirui Hei, Wei Shi وآخرون · 2026

Activation verbalization methods such as Activation Oracle and Natural Language Autoencoders decode hidden representations of large language models into human-readable natural language. However, existing methods can produce incomplete or hallucinated descriptions, making their activation verbalizations difficult to tru …

نسخة أولية وصول مفتوح

RewardExplainer: Learning Reward Model Explanations from Counterfactual Preference Feedback

Jingyi He, Nier Wu, Shuang Liu وآخرون · 2026

Reward models (RMs) are a key component of large language model post-training, providing reward signals for subsequent reinforcement learning. However, conventional discriminative RMs typically output only scalar scores, making it difficult to identify the response behaviors associated with their scoring decisions. Exi …

المؤلفون المشاركون