الباحثون

Jun Xu

المنشورات 4

نسخة أولية وصول مفتوح

What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents

Zhongxiang Sun, Jiahao Yan, Hongkang Zhao وآخرون · 2026

As agents take on long-horizon tasks, users shift from making individual decisions to overseeing autonomous execution. Yet the volume of agent activity and the fragmentation of supporting evidence make it difficult to determine which decisions warrant user verification. We study monitors that identify consequential dec …

نسخة أولية وصول مفتوح

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan وآخرون · 2026

Combining capabilities of multiple expert models trained starting from the same base checkpoint has become increasingly common in frontier language-model post-training. Recent trends suggest that multi-teacher on-policy distillation (MOPD) outperforms conventional off-policy methods. However, despite the higher inferen …

نسخة أولية وصول مفتوح

Solving VeriContest with a Lean-Backed Rust Verifier

Traian Serbanuta, Jun Xu, Andrei Stefanescu وآخرون · 2026

VeriContest is a benchmark of 1007 competitive-programming problems in Rust, each with a Verus specification, a judge-accepted solution, and a Verus proof. Its authors report that proof generation is the bottleneck for frontier models: given the specification and the code, the best model produces an accepted Verus proo …

المؤلفون المشاركون