الباحثون

Xudong Guo

المنشورات 2

نسخة أولية وصول مفتوح

D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?

GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 workloads that measures how effectively agents …

نسخة أولية وصول مفتوح

Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms

Xinjie Shen, Wei Fan, Xudong Guo وآخرون · 2026

Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable outcome signals, and low extension cost. Existing generation pipelines commonly construct an environment before definin …

المؤلفون المشاركون