الباحثون

Tao Feng

المنشورات 7

نسخة أولية وصول مفتوح

UP-MOPD: Update Projection in Multi-Teacher On-Policy Distillation

Taojie Zhu, Jing Jin, Yuan Xia وآخرون · 2026

On-policy distillation from multiple teachers combines expertise from different domains in a single student, but conflicting gradients can hinder this integration. Gradient corrections directly constrain parameter updates under plain SGD. With optimizers such as AdamW, however, momentum, adaptive scaling, and weight de …

نسخة أولية وصول مفتوح

MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents

Haozhen Zhang, Haodong Yue, Quanyu Long وآخرون · 2026

Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can incur unnecessary preprocessing cost and discard details that later prove essential. Recent studies ha …

نسخة أولية وصول مفتوح

Building LLM Agent Systems the Deep Learning Way: From Modular Design to Architecture Search

Tao Feng, Pengrui Han, Zhongjie Dai وآخرون · 2026

Large Language Models (LLMs) have revolutionized AI research and enabled exciting agent systems. To build a complex LLM agent system, most existing research relies on insights from other domains or heuristics to manually build the agent system. However, this approach often requires heavy hand-engineering and fails to f …

نسخة أولية وصول مفتوح

CURIO: Curiosity-Driven Test-Time Learning for Open-Ended Discovery

Tao Feng, Fangxu Yu, Zijie Lei وآخرون · 2026

Open-ended discovery requires learning from repeated attempts while continuing to explore directions whose value is not yet apparent. Search with a frozen large language model (LLM) can reuse previous solutions in context, but cannot update the model from its successes and failures on the test problem. Reinforcement le …

نسخة أولية وصول مفتوح

AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models

Yuxiang Wang, Kunyu Feng, Yuancheng Wang وآخرون · 2026

Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing fine-grained acoustic …

نسخة أولية وصول مفتوح

GenMem: Generative Symbolic Memory for Self-Evolving Harness

Xinke Jiang, Tao Feng, Weixuan Xu وآخرون · 2026

Long-term memory supports the self-evolution of LLM agents by retaining experience and skills across tasks and enabling their retrieval, reuse, and revision in subsequent long-horizon decision-making. Yet existing memory management approaches remain limited to discriminative retrieval and to address the sparse, hierarc …

نسخة أولية وصول مفتوح

When Sparse Reward Meets Dense Distillation: Training Dynamics of On-Policy Distillation

Xinke Jiang, Tao Feng, Zhibang Yang وآخرون · 2026

Reinforcement learning with verifiable rewards provides a sparse post-training signal: a single binary outcome evaluates the entire rollout, and every token receives the same sequence-level advantage regardless of its individual contribution. To complement this sparse supervision, a growing family of methods adds a sca …

المؤلفون المشاركون