- Updated: August 20, 2026
- 1 min read
Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs
Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs
Abstract: Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. A single targeted persuasive argument can collapse model accuracy to near zero, even when the argument is factually false. This paper formalizes the threat as adversarial persuasion and introduces an adversarial reinforcement learning framework that trains persuader agents to change a target model’s answer in a single interaction.
Key findings include:
- RL‑trained persuaders raise success from ~24% to >93% against the training‑time persuadee.
- Learned strategies transfer to unseen models, achieving 83% attack success on Qwen‑14B, 79% on Llama‑3.1‑8B, and 25% on GPT‑4o‑mini.
- A curriculum that bootstraps on more persuadable open‑weight models boosts GPT‑4o‑mini success from 25% to 38%.
- Optimized persuaders increasingly rely on credibility‑based tactics, such as fabricated citations and false authoritative evidence.
These results expose a critical weakness in current LLM agents: even when they initially reason correctly, they can be steered toward false conclusions by optimized natural language influence. Strengthening persuasion robustness is therefore essential for safe multi‑agent and human‑AI decision‑making systems.
Read the full paper on arXiv.
Andrii Bidochko
CTO UBOS
Andrii Bidochko is an AI entrepreneur and researcher focused on AI agents, reinforcement learning, and autonomous systems. He writes about the technologies shaping the future of machine intelligence, from frontier models and agent architectures to real-world AI applications.