Preprint Open access
Does the Model Use the Feature? Separating Steering from Mechanism in LLMs
Internal features in LLMs are often interpreted as mechanisms when they track a concept and their manipulation changes a related behavior. Yet steering can push a feature far outside its natural range, where its effects need not reflect the model's own computation. We examine this inference and propose an empirical con …