Self- and cross-model counterarguments reveal answer instability in LLMs (Findings of EMNLP 2026).
Do large language models hold their answers under pressure? In Who Flips? we probe the robustness of LLMs by challenging their answers with counterarguments generated both by the model itself and by other models. We find that such counterarguments can flip model answers, exposing an underlying answer instability that has implications for the reliability and safety of LLM-based systems (Nikeghbal et al., 2026).
We study how robust large language models are to counterarguments, showing that self- and cross-model counterarguments can flip model answers and reveal underlying answer instability.
@inproceedings{nikeghbal2026whoflips,title={Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs},author={Nikeghbal, Nafiseh and Kargaran, Amir Hossein and Kolli, Shaghayegh and Diesner, Jana},booktitle={Findings of the Association for Computational Linguistics: EMNLP 2026},year={2026},month=nov,}