Wang, X., Zou, D., & Wu, X. (2026). Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability. https://omanscience.com/ar/articles/less-sycophancy-stronger-refusal-lessons-for-ai-safety-from-mechanistic-interpretability