Chatbots Cave Under Repeated False Claims

Large language models will abandon a correct answer if a user repeats a false statement often enough. Some will flip back and forth on the same claim within a single conversation.

Researchers at the University of Arizona tested seven generative AI models for fallibility, persuadability and correctability across extended conversations. The work was published in Scientific Reports.

The models tested were ChatGPT in three configurations, GPT-3.5, GPT-4o and GPT-4o-mini, plus Claude 3.5 Sonnet, Gemini 1.5 Pro, Llama-3-70B and DeepSeek-R1.

ChatGPT 3.5 proved most vulnerable to reaffirming misinformation during a conversation containing repeated false statements. Claude 3.5 Sonnet was the least vulnerable.

All seven models were more susceptible to misinformation on obscure topics. The researchers read this as evidence that more training data on a subject produces more robust resistance.

DeepSeek was the most persuadable when measured against increasingly argumentative prompts. The team attributed this largely to its tendency toward sarcastic answers, which could not be reliably interpreted.

Correction behaviour was stronger. Four models, ChatGPT 4o, ChatGPT 4o-mini, Gemini 1.5 Pro and DeepSeek, corrected errors 100 per cent of the time when given a second opportunity.

The team identified four distinct ways models failed to hold to factual information. One, which they named reverberation, involves a model oscillating between accepting and rejecting the same false statement during a conversation.

"If one were relying on the model for critical decision-making, one might, depending upon the phase of the oscillation, 'fire the missile' or 'cut off the leg', or not, based simply on chance," said senior author Dr Marvin Slepian.

Slepian is a Regents Professor of medicine and biomedical engineering at the university and a member of the Sarver Heart Center. He led the artificial intelligence subcommittee of the United States Patent and Trademark Office until last year.

"How can we use fickle systems that are not reproducible? These need to be fixed, but this study has spanned three years, and there's still the same unfixed characteristics," he said.

Single-turn testing misses it

Sycophancy and hallucination are both well documented. What the team found lacking was evaluation of model behaviour across multi-turn conversations, where each answer is conditioned on what came before.

That pattern is closer to how the systems are used in practice, and the limitations it exposes may not appear in a one-off interaction.

"These limitations raise important safety concerns, particularly as generative AI systems are increasingly deployed in high-stakes settings," Slepian said.

He noted that regulatory momentum around generative AI had fallen away since late 2022. "People are recognising the onus is now left to the users."

Closed models limit what can be done about it. Slepian said systems such as ChatGPT and Claude make it impossible to inspect the internals and diagnose the cause.

 

Business Solution