الملخص
تمت ترجمة أجزاء من هذه الصفحة آلياً وقد تحتوي على أخطاء.
Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, explanations become more useful for predicting model behavior while also increasingly introducing unsupported details and exhibiting writing defects. To assess these changes separately, we introduce a standardized evaluation framework for unstructured NLA explanations, measuring information recoverable from explanations, contextual support for their claims, and writing quality. To address confabulation and writing defects, we move beyond predicting a single activation: explanations can distinguish distributions of possible activations even when their means and optimal point-reconstruction rewards are identical. We introduce Flow-NLA, which models the distribution of activations compatible with an explanation and trains the verbalizer using a diffusion likelihood bound. Across Qwen, Gemma, and Apertus, this richer signal retains the utility gains of point reconstruction while curbing the growth of confabulation and writing defects, opening up a direction for improving activation-derived training to encourage more informative, supported, and readable explanations. Code and evaluation prompts will be made publicly available upon acceptance.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Lek, G., Xia, Z., Chen, P. Y., & Chen, L. Y. (2026). فهم الاختلاق وإعادة التفكير في إعادة البناء في تفسيرات التنشيط. https://omanscience.com/ar/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations
MLA 9
Lek, Gert, et al. "فهم الاختلاق وإعادة التفكير في إعادة البناء في تفسيرات التنشيط." https://omanscience.com/ar/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations.
شيكاغو (المؤلف–التاريخ)
Lek, Gert, Zixuan Xia, Pin-Yu Chen, and Lydia Y. Chen. 2026. "فهم الاختلاق وإعادة التفكير في إعادة البناء في تفسيرات التنشيط." https://omanscience.com/ar/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations.
هارفارد
Lek, G., Xia, Z., Chen, P. Y. and Chen, L. Y. (2026) 'فهم الاختلاق وإعادة التفكير في إعادة البناء في تفسيرات التنشيط', Available at: https://omanscience.com/ar/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations.
فانكوفر
Lek G, Xia Z, Chen PY, Chen LY. فهم الاختلاق وإعادة التفكير في إعادة البناء في تفسيرات التنشيط. https://omanscience.com/ar/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations
IEEE
G. Lek, Z. Xia, P. Y. Chen, and L. Y. Chen, "فهم الاختلاق وإعادة التفكير في إعادة البناء في تفسيرات التنشيط," https://omanscience.com/ar/articles/understanding-confabulation-and-rethinking-reconstruction-in-activation-explanations.