الباحثون

Yutong Dai

المنشورات 4

نسخة أولية وصول مفتوح

CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution

Jixuan Chen, Jiaxin Zhang, Qinyuan Ye وآخرون · 2026

Terminal-agent capability depends jointly on model weights and the runtime harness that formats prompts, binds tools, and handles error recovery. Existing harness-model co-evolution approaches improve both components, yet often treat trajectories produced during harness search as an undifferentiated replay buffer. This …

نسخة أولية وصول مفتوح

Computations of the slice genus and the unknotting number of links via machine learning

Yutong Dai, Oliver Hayman, András Juhász وآخرون · 2026

Links are disjoint unions of circles smoothly embedded in $S^3$. We use reinforcement learning and Bayesian optimisation to obtain new upper bounds on several link invariants that are not known to be algorithmically computable: the slice genus and the unknotting number for links, and the strong slice genus for algebrai …

نسخة أولية وصول مفتوح

CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

Yifan Zhang, Yutong Dai, Viraj Prabhu وآخرون · 2026

Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, while frontier-language-model judges are too expensive to call at every step and cannot be assumed av …

نسخة أولية وصول مفتوح

Opera: A Verbal Critic Framework for Long-horizon Coding Agents

Kai Mei, Zhiyuan Hu, Yutong Dai وآخرون · 2026

Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is delivered. We present …

المؤلفون المشاركون