Who Flips?

Self- and cross-model counterarguments reveal answer instability in LLMs (Findings of EMNLP 2026).

Do large language models hold their answers under pressure? In Who Flips? we probe the robustness of LLMs by challenging their answers with counterarguments generated both by the model itself and by other models. We find that such counterarguments can flip model answers, exposing an underlying answer instability that has implications for the reliability and safety of LLM-based systems (Nikeghbal et al., 2026).

References

2026

  1. Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs
    Nafiseh Nikeghbal, Amir Hossein Kargaran, Shaghayegh Kolli, and 1 more author
    In Findings of the Association for Computational Linguistics: EMNLP 2026, Nov 2026