Preprint Open access
Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development
Parallel coding agents can produce patches that work alone but fail when merged. This happens when one agent changes an interface or rule that another agent still relies on. We study these failures with stale, a benchmark for semantic coordination. Our evaluation runs the same tests on each patch alone and on their com …