الملخص

تمت ترجمة أجزاء من هذه الصفحة آلياً وقد تحتوي على أخطاء.

In hospital workflows, electronic health records (EHRs) are often noisy, and may not contain the evidence needed to confirm events or measurements referenced in a clinical query. Even when database retrieval succeeds, clinical agents can overlook such discrepancies and return plausible but unsupported answers. We introduce EHR-RobustGym, a scalable and interactive environment for evaluating and training robust clinical agents grounded in noisy EHRs. Built on MIMIC-IV hospital records (365K patients, 31 tables, and over 500M records), EHR-RobustGym comprises 5,486 Clean-Noise pairs spanning six clinical intents and both patient-level and population-level queries. The pairs test robustness to Record-level, Value-level, and Query-level noise, while interactive SQL/Python execution and outcome verification support trajectory collection and training. Evaluating multiple LLMs reveals substantial robustness gaps: average task success across proprietary and large-scale open-weight models drops from 62.2% on Clean questions to 37.9% on Noise questions. At k=4, pass^k consistency falls below 50% for most evaluated models, exposing instability in clinical task completion. Supervised fine-tuning and reinforcement learning in EHR-RobustGym improve performance, with gains generalizing to five external EHR benchmarks. Together, these results position EHR-RobustGym as a testbed for evaluating and improving the evidence-grounded robustness of clinical agents.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Qiao, Y., Jin, Y., Liu, L., Shen, Y., Wang, J., Gu, J., & Chu, Z. (2026). EHR-RobustGym: قياس وتدريب الوكلاء على الاستدلال السريري المتين. https://omanscience.com/ar/articles/ehr-robustgym-benchmarking-and-training-agents-for-robust-clinical-reasoning

MLA 9

Qiao, Yitong, et al. "EHR-RobustGym: قياس وتدريب الوكلاء على الاستدلال السريري المتين." https://omanscience.com/ar/articles/ehr-robustgym-benchmarking-and-training-agents-for-robust-clinical-reasoning.

شيكاغو (المؤلف–التاريخ)

Qiao, Yitong, Yancheng Jin, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, and Zhixuan Chu. 2026. "EHR-RobustGym: قياس وتدريب الوكلاء على الاستدلال السريري المتين." https://omanscience.com/ar/articles/ehr-robustgym-benchmarking-and-training-agents-for-robust-clinical-reasoning.

هارفارد

Qiao, Y., Jin, Y., Liu, L., Shen, Y., Wang, J., Gu, J. and Chu, Z. (2026) 'EHR-RobustGym: قياس وتدريب الوكلاء على الاستدلال السريري المتين', Available at: https://omanscience.com/ar/articles/ehr-robustgym-benchmarking-and-training-agents-for-robust-clinical-reasoning.

فانكوفر

Qiao Y, Jin Y, Liu L, Shen Y, Wang J, Gu J, et al. EHR-RobustGym: قياس وتدريب الوكلاء على الاستدلال السريري المتين. https://omanscience.com/ar/articles/ehr-robustgym-benchmarking-and-training-agents-for-robust-clinical-reasoning

IEEE

Y. Qiao, Y. Jin, L. Liu, Y. Shen, J. Wang, J. Gu, and Z. Chu, "EHR-RobustGym: قياس وتدريب الوكلاء على الاستدلال السريري المتين," https://omanscience.com/ar/articles/ehr-robustgym-benchmarking-and-training-agents-for-robust-clinical-reasoning.