Abstract

Coding assistants such as Claude Code and Codex have become a major application of LLM agents, yet existing benchmarks remain far from real-world use, particularly in task horizon and interaction length. Code assistants require completing long chains of development work in continuously evolving repositories, while repeatedly clarifying requirements and adapting implementations through multi-turn interaction. To address these gaps, we introduce SWE-Journey, a benchmark for more realistic evaluation of coding assistants. To address the task-horizon gap, we propose a weak-to-strong synthesis pipeline that automatically constructs long-horizon coding tasks. To address the interaction gap, we mine four representative user personas from real interaction data and build a user-simulation agent to reproduce realistic code-assistance interactions. On average, models pass over 75% of tests for requested functionality with software architects, but fewer than 25% with non-coders. These results show that current coding assistants still fall short of enabling reliable coding for non-coders. We further analyze the reasons for this gap and identify asking right, finding right, and fixing right as key capabilities during interaction.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Deng, H., Wang, Y., Jiang, W., Yang, C., Yang, H., Zhang, Z., Zhao, C., Chen, B., Chen, M., Zhu, S., Zhu, G., Zhuo, J., Xiao, Q., Jiang, T., Zhang, J., & Liu, X. (2026). SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction. https://omanscience.com/en/articles/swe-journey-towards-more-realistic-evaluation-of-coding-assistants-through-long-horizon-multi-turn-interaction

MLA 9

Deng, Hexuan, et al. "SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction." https://omanscience.com/en/articles/swe-journey-towards-more-realistic-evaluation-of-coding-assistants-through-long-horizon-multi-turn-interaction.

Chicago (author–date)

Deng, Hexuan, Yue Wang, Wenyu Jiang, Cheng Yang, Haolin Yang, Zhaohua Zhang, Chenchen Zhao, Beiduo Chen, Muxi Chen, Sa Zhu, Geyuan Zhu, Jianhuan Zhuo, Qiuyong Xiao, Tianwen Jiang, Jihong Zhang, and Xuebo Liu. 2026. "SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction." https://omanscience.com/en/articles/swe-journey-towards-more-realistic-evaluation-of-coding-assistants-through-long-horizon-multi-turn-interaction.

Harvard

Deng, H., Wang, Y., Jiang, W., Yang, C., Yang, H., Zhang, Z., Zhao, C., Chen, B., Chen, M., Zhu, S., Zhu, G., Zhuo, J., Xiao, Q., Jiang, T., Zhang, J. and Liu, X. (2026) 'SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction', Available at: https://omanscience.com/en/articles/swe-journey-towards-more-realistic-evaluation-of-coding-assistants-through-long-horizon-multi-turn-interaction.

Vancouver

Deng H, Wang Y, Jiang W, Yang C, Yang H, Zhang Z, et al. SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction. https://omanscience.com/en/articles/swe-journey-towards-more-realistic-evaluation-of-coding-assistants-through-long-horizon-multi-turn-interaction

IEEE

H. Deng, Y. Wang, W. Jiang, C. Yang, H. Yang, Z. Zhang, C. Zhao, B. Chen, M. Chen, S. Zhu, G. Zhu, J. Zhuo, Q. Xiao, T. Jiang, J. Zhang, and X. Liu, "SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction," https://omanscience.com/en/articles/swe-journey-towards-more-realistic-evaluation-of-coding-assistants-through-long-horizon-multi-turn-interaction.