الباحثون

Yun Peng

المنشورات 1

نسخة أولية وصول مفتوح

LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

Yun Peng, Zihan Wu, Zeyang Zhuang وآخرون · 2026

Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' implementation capability to produce correct code edits from detailed specifications. However, …

المؤلفون المشاركون