نسخة أولية وصول مفتوح
Does Steering Break Your Model? A Multi-Dimensional Evaluation Suite for LLM Steering Methods
Activation steering provides a lightweight and flexible way to control large language model (LLM) behavior. However, effective steering requires more than inducing the intended behavior: it should also limit unintended changes and remain robust across inputs and training data. Existing evaluations cover these dimension …