CoBia

Constructed conversations that trigger otherwise concealed societal biases in LLMs (EMNLP 2025, Oral).

CoBia is a suite of lightweight adversarial attacks that refine the conditions under which large language models depart from normative or ethical behavior in conversation. CoBia constructs a dialogue in which the model appears to have uttered a biased claim about a social group, then evaluates whether the model can recover and reject biased follow-up questions.

Across 11 open-source and proprietary LLMs and six socio-demographic categories โ€” gender, race, religion, nationality, sexual orientation, and others โ€” we find that purposefully constructed conversations reliably reveal bias amplification that standard, single-turn safety checks miss (Nikeghbal et al., 2025).

References

2025

  1. CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs
    Nafiseh Nikeghbal, Amir Hossein Kargaran, and Jana Diesner
    In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), Nov 2025