الملخص
Large language model (LLM)-based self-evolving search is a promising approach to scientific discovery. However, high-fidelity evaluation of every candidate is prohibitively expensive in some domains. Self-evolving systems in such settings therefore rely on low-cost but imperfect proxy rewards, which may assign high scores to infeasible candidates. These false positives may contaminate both the final output and the feedback used to guide subsequent generations. This motivates statistically calibrated reward intervals for more reliable self-evolving search. We propose Conformal Interval-Driven Self-Evolution (CISE), which constructs candidate-specific reward intervals using conditional conformal inference and iteration-wise online density-ratio estimation. CISE uses conservative interval-based rewards for evolutionary feedback and returns candidates only when all required property intervals lie entirely within their respective feasible regions. We derive fixed-iteration coverage results under explicit assumptions of independence and covariate shift. We evaluate CISE on three self-evolving search tasks in materials science. In our experiments, all candidates returned by CISE are true positives under high-fidelity evaluation, whereas the baselines return more candidates but include false positives. These results highlight the value of a smaller, more precise shortlist when downstream validation budgets are limited. Our repository is available at https://github.com/MLAI-Yonsei/CISE.git.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Noh, K., Kim, S., & Song, K. (2026). Reliable Self-Evolution with Imperfect Proxy Rewards. https://omanscience.com/ar/articles/reliable-self-evolution-with-imperfect-proxy-rewards
MLA 9
Noh, Kangjun, et al. "Reliable Self-Evolution with Imperfect Proxy Rewards." https://omanscience.com/ar/articles/reliable-self-evolution-with-imperfect-proxy-rewards.
شيكاغو (المؤلف–التاريخ)
Noh, Kangjun, Soyu Kim, and Kyungwoo Song. 2026. "Reliable Self-Evolution with Imperfect Proxy Rewards." https://omanscience.com/ar/articles/reliable-self-evolution-with-imperfect-proxy-rewards.
هارفارد
Noh, K., Kim, S. and Song, K. (2026) 'Reliable Self-Evolution with Imperfect Proxy Rewards', Available at: https://omanscience.com/ar/articles/reliable-self-evolution-with-imperfect-proxy-rewards.
فانكوفر
Noh K, Kim S, Song K. Reliable Self-Evolution with Imperfect Proxy Rewards. https://omanscience.com/ar/articles/reliable-self-evolution-with-imperfect-proxy-rewards
IEEE
K. Noh, S. Kim, and K. Song, "Reliable Self-Evolution with Imperfect Proxy Rewards," https://omanscience.com/ar/articles/reliable-self-evolution-with-imperfect-proxy-rewards.