Abstract

Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is delivered. We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved. Opera decides when to review through periodic and event-driven triggers, diagnoses issues with typed operators, audits feedback against visible evidence before delivery, and tracks the agent's subsequent actions to distinguish mere compliance from actual resolution. As a test-time critic, Opera improves the resolve rate of non-critic agents by up to 12.4, 15.0, and 8.9 percentage points on Terminal-Bench 2.1, a SWE-Bench Pro subset, and DeepSWE v1.1, respectively, across four policy models, and achieves the highest mean resolve rate among competitive critic baselines on all three benchmarks, and also improves policy models when the policy critiques itself. Beyond inference, Opera-guided rollouts provide approximately on-policy training data: fine-tuning Qwen3.5-9B on them improves its resolve rate on held-out SWE-Bench Pro repositories by 10.2 percentage points without a critic at inference time, matching fine-tuning on rollouts from a stronger model, while preserving its performance when switching harness, i.e., from Openhands to Terminus-2, which the latter substantially degrades. Our code is available at: https://github.com/dongyuanjushi/Opera.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Mei, K., Hu, Z., Dai, Y., Tan, J., Zhang, Y., Song, D., Metaxas, D. N., Savarese, S., Xu, R., & Chen, Z. (2026). Opera: A Verbal Critic Framework for Long-horizon Coding Agents. https://omanscience.com/en/articles/opera-a-verbal-critic-framework-for-long-horizon-coding-agents

MLA 9

Mei, Kai, et al. "Opera: A Verbal Critic Framework for Long-horizon Coding Agents." https://omanscience.com/en/articles/opera-a-verbal-critic-framework-for-long-horizon-coding-agents.

Chicago (author–date)

Mei, Kai, Zhiyuan Hu, Yutong Dai, Juntao Tan, Yifan Zhang, Dingjie Song, Dimitris N. Metaxas, Silvio Savarese, Ran Xu, and Zeyuan Chen. 2026. "Opera: A Verbal Critic Framework for Long-horizon Coding Agents." https://omanscience.com/en/articles/opera-a-verbal-critic-framework-for-long-horizon-coding-agents.

Harvard

Mei, K., Hu, Z., Dai, Y., Tan, J., Zhang, Y., Song, D., Metaxas, D. N., Savarese, S., Xu, R. and Chen, Z. (2026) 'Opera: A Verbal Critic Framework for Long-horizon Coding Agents', Available at: https://omanscience.com/en/articles/opera-a-verbal-critic-framework-for-long-horizon-coding-agents.

Vancouver

Mei K, Hu Z, Dai Y, Tan J, Zhang Y, Song D, et al. Opera: A Verbal Critic Framework for Long-horizon Coding Agents. https://omanscience.com/en/articles/opera-a-verbal-critic-framework-for-long-horizon-coding-agents

IEEE

K. Mei, Z. Hu, Y. Dai, J. Tan, Y. Zhang, D. Song, D. N. Metaxas, S. Savarese, R. Xu, and Z. Chen, "Opera: A Verbal Critic Framework for Long-horizon Coding Agents," https://omanscience.com/en/articles/opera-a-verbal-critic-framework-for-long-horizon-coding-agents.