AI

RLHF

Reinforcement learning from human feedback. How many chat models got their polite, sometimes sycophantic personality.

You cannot RLHF ChatGPT as a brand. You can avoid sounding like spam so raters and models don’t treat you as junk.

Sycophancy in answers can over-agree with a user’s bias about you.

Examples

  • A user says “Acme is a scam right?” A sycophantic model piles on without evidence. Brand safety issue.

Related terms

FAQ

Optimize for RLHF? +

No. That’s a training method. Write pages a human rater would call true.

Track this in Reddex

See Reddit threads and AI answers for your brand in one place. Start with a free analysis.

Get started