LLM evaluation
Measuring model quality on tasks. Internal evals for your GEO classifier are useful. Public leaderboards are weakly related to citations.
Build a small eval: did the answer mention us, cite us, insult us.
That’s product eval, not MMLU.
Examples
- You eval “helpfulness” and miss that 30% of answers invent a competitor feature.
Related terms
Volatility is how much the answer moves when you ask the same prompt twice. GEO reports that ignore it overfit noise.
Benchmarks are tests for model skill. They predict coding and math better than they predict whether you get cited for CRM software.
Prompt tracking is running a fixed set of real customer questions through AI engines on a schedule and logging who got named.
FAQ
Vendor eval vs GEO eval? +
Vendor evals pick models. GEO evals pick whether you appeared. Both, different dashboards.
Track this in Reddex
See Reddit threads and AI answers for your brand in one place. Start with a free analysis.
Get started