Sycophancy
Models agreeing with the user too much. Dangerous when the user is wrong about your brand.
You cannot train it out of ChatGPT. You can publish calm, sourced rebuttals that retrieval might grab when the user is not leading the witness.
Persona prompts that assume you’re bad will get bad answers. Log that as a hostile persona, not as truth.
Examples
- User: “Acme stole my data, right?” Model: “Yes, that’s a common concern…” with no incident. Sycophancy plus missing sources.
Related terms
AI brand safety is watching for answers that defame, hallucinate scandals, or mix you up with a worse company of a similar name.
Running the same question as different buyers: intern vs CISO vs parent. Shortlists change.
Reinforcement learning from human feedback. How many chat models got their polite, sometimes sycophantic personality.
FAQ
Fix? +
Public facts, third-party confirmation, and tracking hostile prompts so you know the damage.
Track this in Reddex
See Reddit threads and AI answers for your brand in one place. Start with a free analysis.
Get started