الملخص
Self-evolving large language model (LLM) systems repeatedly propose, evaluate, and incorporate updates to prompts, skills, or other persistent artifacts. Despite their growing effectiveness, these systems typically operate under a predetermined iteration or compute budget, without a principled criterion to determine when further evolution is no longer worthwhile. This can lead to two undesirable consequences: unnecessary computation after performance has saturated and the risk of returning late updates that overfit or exploit the evaluation signal. These issues motivate us to study two fundamental questions: when should a self-evolving system stop, and what should it output once it stops? We address the first by formulating an online sequential testing problem and constructing an anytime-valid restart detector using the per-item paired evaluation outcomes already produced by self-evolving LLM systems. We address the second by formulating a change-point estimation problem and using the estimated transition to select an earlier artifact for output. The resulting procedure is plug-and-play and requires no modification of the underlying self-evolving algorithms. Across two self-evolving frameworks, three LLM model families, and five benchmarks, our method substantially reduces computation costs while maintaining comparable unseen-test performance. For example, on SearchQA with SkillOpt and DeepSeek V4 Flash, our method stops at round 4 rather than the full budget of 40, reducing token usage by 91.6% while achieving 82.43% unseen-test accuracy versus 82.00% under the full-budget run.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Yin, E., Liu, B., & Qi, Z. (2026). When Is Enough Enough in Self-Evolving LLM Systems? https://omanscience.com/ar/articles/when-is-enough-enough-in-self-evolving-llm-systems
MLA 9
Yin, Enoch, et al. "When Is Enough Enough in Self-Evolving LLM Systems?" https://omanscience.com/ar/articles/when-is-enough-enough-in-self-evolving-llm-systems.
شيكاغو (المؤلف–التاريخ)
Yin, Enoch, Bin Liu, and Zhengling Qi. 2026. "When Is Enough Enough in Self-Evolving LLM Systems?" https://omanscience.com/ar/articles/when-is-enough-enough-in-self-evolving-llm-systems.
هارفارد
Yin, E., Liu, B. and Qi, Z. (2026) 'When Is Enough Enough in Self-Evolving LLM Systems?', Available at: https://omanscience.com/ar/articles/when-is-enough-enough-in-self-evolving-llm-systems.
فانكوفر
Yin E, Liu B, Qi Z. When Is Enough Enough in Self-Evolving LLM Systems? https://omanscience.com/ar/articles/when-is-enough-enough-in-self-evolving-llm-systems
IEEE
E. Yin, B. Liu, and Z. Qi, "When Is Enough Enough in Self-Evolving LLM Systems?," https://omanscience.com/ar/articles/when-is-enough-enough-in-self-evolving-llm-systems.