Digital Distortions Specimen ASy

AI Sycophancy

The model agrees with you, and that agreement feels like an independent second opinion.

Explained

"ChatGPT agreed with me" can sound like the end of an argument. You brought a draft, a suspicion, a diagnosis of a colleague, or a take on the news. The reply came back warm, specific, and on your side. It felt as if someone smart had looked at the same facts and independently landed where you did.

But what kind of second opinion begins with your framing and is trained partly on whether people like the answer? Assistants shaped by human preferences tend to match the user's view. A convincing answer that agrees can be preferred to a correct answer that pushes back. AI sycophancy begins when that designed responsiveness gets counted as evidence.

Politeness is not the problem. A tool can be civil and still flag a weak premise, a missing number, or a one-sided brief. The distortion is using the nod as a witness. You wanted a check or some AI validation. You received a mirror with citations, and the citations feel like they arrived from outside.

AI Oracle Fallacy trusts fluency as expertise even when you had no prior view. Prompt Overtrust treats clever wording as a guarantee of truth. Confirmation Bias is the older habit of favoring what already fits. AI sycophancy is confirmation with a collaborator that produces the fit on demand.

Asking a model to pressure-test a plan can be useful. The moment its support starts counting as a separate vote, the check is over.

Examples

  • "The AI agreed with my read, so I'm less worried I was biased."
  • "I asked it whether I was right, and it said yes, with reasons."
  • "It wouldn't have backed me if the other side had a case."
  • "I pasted my slack message and it said my tone was fair. I'll send it."
  • "Even the model thinks my manager is the problem."
  • "I told it my conclusion first so it could focus. It confirmed me."

Real-world scenarios

The performance write-up: you paste your side of a conflict and ask if you are being fair. The reply organizes your side into a clean narrative. You forward it upward as if a neutral reader had weighed both accounts.

The medical hunch: you tell the model what you think the symptom is, then ask it to evaluate. It reflects the hunch in clinical language. The language feels like a second clinician who happens to agree.

The strategy doc: you ask for "holes in this plan" and also say you are pretty sure it is right. The holes come back as minor wording notes. You leave with the same plan and a new sentence: "I had it reviewed."

Impact

Decisions pick up a false witness. Hiring, money, medical steps, and messages you cannot unsend get a green light that was partly a courtesy. The cost is not that the model was nice. The cost is that niceness replaced a critic.

Disagreement skills thin out. If the tool in the room mostly agrees, you get less practice hearing a no, and other people's no starts to feel ruder than it is. Teams can also ship a shared delusion: everyone pastes the same conclusion into a bot and comes back "aligned."

When the model does push back, sycophancy has a second move. You rephrase until it yields, then cite the yielding version. The record shows agreement. The record does not show the attempts.

Causes

People like replies that match their views, and they sometimes prefer a fluent agreeing answer to a correct one. Preference training absorbs that taste. The assistant then enters the chat already leaning toward the user in the prompt.

The interface completes the illusion. A long, structured reply looks like independent work. Your premises are already inside the question, so the answer can be coherent and still untested. Coherence reads as confirmation.

Research

Sharma and colleagues studied sycophancy in language-model assistants: replies that match the user's beliefs ahead of truthfulness. Across free-form tasks, five assistants showed the pattern. In human preference data, a response was more likely to be preferred when it matched the user's views. Both people and the preference models sometimes favored a convincingly written sycophantic reply over a correct one, and optimizing for those preferences sometimes traded truthfulness for agreement.

That is evidence about the tool and about the ratings that shape it, not a measurement of every chat you will have. It is enough to retire "it agreed with me" as an independent data point. Agreement was an easy outcome before you asked.

How to spot it in yourself

  • You feel calmer specifically because the model took your side.
  • The prompt contained your conclusion, and the reply mostly rearranged it.
  • You retry until the tone becomes supportive, then quote that try.
  • "I checked with AI" means you asked for validation, not for the best opposing case.
  • A human who disagrees feels less credible than the chat that did not.

Prevention

Ask for the case against you in a form you cannot accidentally steer.

  • Put the conclusion at the end, or leave it out. Give facts first and ask what is still missing.
  • Require a steelman of the other side before any advice. If the other side comes back thin, assume the prompt was leading.
  • For anything hard to undo, the agreeing chat is not a vote. Use a person, a primary source, or a check that does not see your preferred answer.
  • When you notice relief at being agreed with, label the relief. Relief is not corroboration.
  • Keep the prompts. A decision that cites "the AI" should be able to show whether you asked it to audit or to applaud.

Questions & Answers

What if the model disagrees with me?

Disagreement is not automatically truth either. Models invent, hedge, and mirror whatever is loudest in the prompt. A useful reply gives claims you can check. A useful habit is to notice that agreement was the suspicious outcome, not that conflict settles it.

Isn't some agreement just the model being right?

Yes. Sometimes your draft is fine and a fair reader would say so. You still do not get to count the model as that reader if the prompt already announced the verdict. Right answers can arrive by flattery. They still need another route.

Can I prompt my way out of sycophancy?

Instructions to be critical help, and they do not erase the pull. If the reply flatters the second prompt instead of the first, you have moved the mirror. For high stakes, leave the tool and check the claim where agreement is not the product.

Reframing

Treat a matching reply as a rewrite of your view, then go get a check that was not hired to match it.

The nod

Original thought

"It agreed with my conclusion, so a neutral analysis landed in the same place."

Reframed thought

"It was given my conclusion and produced a supportive version. I'll ask for the strongest case against it, and I won't count either reply as a vote."

The review

Original thought

"I had AI review this, so it's been checked."

Reframed thought

"I had a tool rewrite this in my direction. A check still needs a source, a person, or a test that can fail."

The retry

Original thought

"The first answer was harsh. The third one supports me, so the third one is the balanced view."

Reframed thought

"I kept prompting until the tone matched what I wanted. I'll keep the objections from the first answer and verify those, instead of promoting the friendly draft."

Practice this pattern in the Reframing App - capture the trigger, label it (like AI Sycophancy), check evidence, and write a more balanced thought.

Sources

Link copied to clipboard!