الباحثون

Yu Cheng

المنشورات 10

نسخة أولية وصول مفتوح

Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts

Yunkai Chai, Tong Zhu, Xiaoye Qu وآخرون · 2026

Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token. However, traditional training paradigms enforce a static choice of top-$k$ experts, which converts a continuous routing distribution into a rigid step function. This constraint introduces a brittle …

نسخة أولية وصول مفتوح

MIRA: A Musical Intent Refinement Agent for Aligning Text-to-Music Generation with User Intent

Zekai Liu, Zhilin Wang, Xuzheng He وآخرون · 2026

Text-to-music systems produce increasingly convincing audio, yet evaluation reveals little about whether the result matches user intent. A global text-audio relevance score can overlook the implicit intent in underspecified prompts and mask failures in specific requirements, such as instrumentation, structure, rhythm, …

نسخة أولية وصول مفتوح

An Empirical Study of Agent Skills' Downstream Utility

Yu Cheng, Dehai Zhao, Zhongxin Liu وآخرون · 2026

Agent Skills package procedural guidance and resources for reuse, but a relevant Skill does not necessarily improve task performance. Existing studies characterize Skill content and evaluate downstream performance, yet provide limited explanations of how utility depends on content, execution configuration, and multi-Sk …

نسخة أولية وصول مفتوح

Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation

Wenxuan Wang, Zekai Liu, Weinan Zhang وآخرون · 2026

Wrapping an image generation model in an agentic harness can effectively boost Text-to-Image task performance: the harness can leverage memory, skills, workflow orchestration, result verification, and iterative refinement to continually construct and revise prompts, thereby eliciting better images. These gains, however …

نسخة أولية وصول مفتوح

SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning

Zihao Chen, Fanxiang Xiong, Hongran Ren وآخرون · 2026

Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count. Un …

نسخة أولية وصول مفتوح

SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time

Yu Cheng, Yongkang Hu, Shuaijie Ma وآخرون · 2026

LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches commonly rely on multiple rounds of optimization over fixed and repeatedly accessible task d …

نسخة أولية وصول مفتوح

TMCS: Tool-Grounded Multi-Agent Reasoning for Compositional Chemical Problem Solving

Shengqin Wang, Jie Jin, Yu Cheng وآخرون · 2026

Despite the promise of Large Language Models (LLMs) in computational chemistry, rigorous combinatorial chemistry problems remain difficult because they require quantitatively constrained molecular modification, candidate validation, and systematic revision after failed attempts. Existing tool-augmented chemical agents …

نسخة أولية وصول مفتوح

The Past Frames the Future: Memory for Autoregressive Video Generation

Harold Haodong Chen, Rongjin Guo, Disen Lan وآخرون · 2026

Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the generated sequence expands, …

نسخة أولية وصول مفتوح

HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing

Jianyu Wei, Yizhao Gao, Qihao Zhang وآخرون · 2026

Long-horizon and multi-turn agents typically generate short actions and process long observations from tools and environments. This growing context demands efficient prefill, compact KV-cache storage, and accurate long-context retrieval. To meet these demands, we introduce HySparse2, a hybrid sparse attention architect …

المؤلفون المشاركون