الباحثون

Thomas Fel

المنشورات 5

نسخة أولية وصول مفتوح

Language Model Activations Inhabit Privileged Error-Correcting Basins

Language models exhibit remarkable robustness, continuing to produce coherent text even when their activations are perturbed by interventions like linear steering. We hypothesize that this robustness is a result of passive dynamics, i.e., constraining mechanisms in the forward pass that funnel activations toward "good" …

نسخة أولية وصول مفتوح

When Models Don't Manipulate Manifolds: The Geometry of a Comparison Task

One of the current premises of mechanistic interpretability research is that detailed accounts of the geometry of neural network representations can tell us how models perform computations, and how to effectively intervene on them. While low dimensional manifolds have been observed for multiple concepts in the literatu …

نسخة أولية وصول مفتوح

Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations

Leon Bergen, Usha Bhalla, Andrew Lee وآخرون · 2026

As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential. Does it leave a telltale signature in model representations? This work analyzes how reward hacking is represented internally in frontier open source LLMs, and how those representations can be used to understand and discov …

المؤلفون المشاركون