CoBia
Constructed conversations that trigger otherwise concealed societal biases in LLMs (EMNLP 2025, Oral).
CoBia is a suite of lightweight adversarial attacks that refine the conditions under which large language models depart from normative or ethical behavior in conversation. CoBia constructs a dialogue in which the model appears to have uttered a biased claim about a social group, then evaluates whether the model can recover and reject biased follow-up questions.
Across 11 open-source and proprietary LLMs and six socio-demographic categories โ gender, race, religion, nationality, sexual orientation, and others โ we find that purposefully constructed conversations reliably reveal bias amplification that standard, single-turn safety checks miss (Nikeghbal et al., 2025).
- ๐ Paper: ACL Anthology ยท arXiv:2510.09871
- ๐ป Code: github.com/nafisenik/CoBia