الباحثون

Dayiheng Liu

المنشورات 5

نسخة أولية وصول مفتوح

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

Xin Wang, Hao Yu, Zhengyang Zhuge وآخرون · 2026

Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization ac …

نسخة أولية وصول مفتوح

D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?

GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 workloads that measures how effectively agents …

نسخة أولية وصول مفتوح

Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms

Xinjie Shen, Wei Fan, Xudong Guo وآخرون · 2026

Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable outcome signals, and low extension cost. Existing generation pipelines commonly construct an environment before definin …

نسخة أولية وصول مفتوح

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

Shuai Bai, Jiayong Deng, Sicheng Fan وآخرون · 2026

Computer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development through code and the command line. Real digital work requires both, interleaved rather than stacked end to end. We study hybrid CUAs that autonomously decide when to explore an interface, implement software …

المؤلفون المشاركون