الملخص

تمت ترجمة أجزاء من هذه الصفحة آلياً وقد تحتوي على أخطاء.

Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, explanations become more useful for predicting model behavior while also increasingly introducing unsupported details and exhibiting writing defects. To assess these changes separately, we introduce a standardized evaluation framework for unstructured NLA explanations, measuring information recoverable from explanations, contextual support for their claims, and writing quality. To address confabulation and writing defects, we move beyond predicting a single activation: explanations can distinguish distributions of possible activations even when their means and optimal point-reconstruction rewards are identical. We introduce Flow-NLA, which models the distribution of activations compatible with an explanation and trains the verbalizer using a diffusion likelihood bound. Across Qwen, Gemma, and Apertus, this richer signal retains the utility gains of point reconstruction while curbing the growth of confabulation and writing defects, opening up a direction for improving activation-derived training to encourage more informative, supported, and readable explanations. Code and evaluation prompts will be made publicly available upon acceptance.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Lek, G., Xia, Z., Chen, P. Y., & Chen, L. Y. (2026). فهم الاختلاق وإعادة التفكير في إعادة البناء في تفسيرات التنشيط. https://omanscience.com/ar/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations

MLA 9

Lek, Gert, et al. "فهم الاختلاق وإعادة التفكير في إعادة البناء في تفسيرات التنشيط." https://omanscience.com/ar/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations.

شيكاغو (المؤلف–التاريخ)

Lek, Gert, Zixuan Xia, Pin-Yu Chen, and Lydia Y. Chen. 2026. "فهم الاختلاق وإعادة التفكير في إعادة البناء في تفسيرات التنشيط." https://omanscience.com/ar/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations.

هارفارد

Lek, G., Xia, Z., Chen, P. Y. and Chen, L. Y. (2026) 'فهم الاختلاق وإعادة التفكير في إعادة البناء في تفسيرات التنشيط', Available at: https://omanscience.com/ar/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations.

فانكوفر

Lek G, Xia Z, Chen PY, Chen LY. فهم الاختلاق وإعادة التفكير في إعادة البناء في تفسيرات التنشيط. https://omanscience.com/ar/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations

IEEE

G. Lek, Z. Xia, P. Y. Chen, and L. Y. Chen, "فهم الاختلاق وإعادة التفكير في إعادة البناء في تفسيرات التنشيط," https://omanscience.com/ar/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations.