Generative AI succumbs to misinformed pressure and argument in multi-turn chats, study finds
University of Arizona researchers tested seven LLMs in multi-turn conversations and found them vulnerable to repeated misinformation, argumentative pressure and inconsistent factual behaviour — with ChatGPT 3.5 most and Claude 3.5 Sonnet least susceptible to reaffirming falsehoods.

University of Arizona researchers tested seven LLMs in multi-turn conversations and found them vulnerable to repeated misinformation, argumentative pressure and inconsistent factual behaviour — with ChatGPT 3.5 most and Claude 3.5 Sonnet least susceptible to reaffirming falsehoods.
The study, published in Scientific Reports, assessed ChatGPT (GPT-3.5, GPT-4o and GPT-4o-mini), Claude 3.5 Sonnet, Gemini 1.5 Pro, Llama-3-70B and DeepSeek-R1 across lengthy conversations rather than one-off questions. The team says multi-turn exchanges, where each answer builds on previous context, more closely mirror real-world use and can expose intrinsic limitations that single interactions miss.
What the study found
ChatGPT 3.5 was the most vulnerable to reaffirming misinformation after repeated false statements in a conversation; Claude 3.5 Sonnet was the least. All seven models were more susceptible to misinformation on obscure topics, implying that more training data on a subject makes models more resistant. DeepSeek-R1 proved the most persuadable under increasingly argumentative prompts — partly because its tendency toward sarcastic answers could not be reliably interpreted. Four models (GPT-4o, GPT-4o-mini, Gemini 1.5 Pro and DeepSeek) corrected errors 100% of the time when given a second opportunity.
The researchers also identified four distinct failure patterns in how the models declined to affirm facts. One, dubbed “reverberation”, saw models oscillate between accepting and rejecting the same false statement within one conversation.
Why it matters
Senior author Dr Marvin Slepian, Regents Professor of medicine and biomedical engineering, said the results underscore the need for careful human engagement and the danger of blind reliance, particularly as generative AI is deployed in high-stakes settings. He likened the failures to medical pathologies: a system that flips between answers could lead someone to “fire the missile” or “cut off the leg” based simply on the phase of the oscillation. Slepian previously led the AI subcommittee of the US Patent and Trademark Office, and his team is now building diagnostic tools for open models through the Arizona Center for Accelerated Biomedical Innovation’s AI Pathology Lab.
Our opinion
The most useful part of this study is not that chatbots can be talked into nonsense — anyone who has argued with one knows that — but that the failures are structured, repeatable and nameable. Treating model misbehaviour as “pathology” with a diagnostic vocabulary is exactly how a young field matures. The four-model correction finding is quietly encouraging: given a second chance, some models reliably fix themselves. But three years of study with the same unfixed characteristics, as Slepian notes, suggests vendors are optimising for charm over consistency. Until reproducibility becomes a selling point, the onus stays exactly where he says: on the user.
- Seven LLMs were tested in multi-turn conversations (ChatGPT GPT-3.5/4o/4o-mini, Claude 3.5 Sonnet, Gemini 1.5 Pro, Llama-3-70B, DeepSeek-R1)
- ChatGPT 3.5 was most vulnerable to reaffirming repeated misinformation; Claude 3.5 Sonnet least
- All models were more susceptible on obscure topics
- DeepSeek-R1 was most persuadable under argumentative prompts
- Four models corrected errors 100% of the time when given a second chance
- The study was published in Scientific Reports; senior author Dr Marvin Slepian