Preprint Open access
Acmite: Mitigating Gender Bias in LLMs through Concept-Guided Mutual Information
Large language models (LLMs) can reproduce social stereotypes from their training data, motivating extensive research on model debiasing. However, existing methods often rely on explicit biased examples or predefined group-term substitutions, making them sensitive to wording and less effective at capturing stereotype c …