AI inference
Inference is running the trained model to produce an answer. Every ChatGPT reply is inference. It costs money, which is why vendors cache and skip retrieval.
If they cache, your fresh page waits. Volatility and staleness both come from inference economics, not mysticism.
Examples
- A popular prompt returns a cached roundup for hours. Your embargoed launch is invisible until the cache dies.
Related terms
Volatility is how much the answer moves when you ask the same prompt twice. GEO reports that ignore it overfit noise.
Models that spend more inference compute to think longer (o-series, R1, etc.). They can pick sources more carefully — or overthink a Reddit anecdote.
Spending more compute at inference to think longer. Reasoning models. Latency and cost go up; answers may get more careful.
FAQ
Can I pay to invalidate cache? +
Not as a random brand. Publish, get others to talk, wait.
Track this in Reddex
See Reddit threads and AI answers for your brand in one place. Start with a free analysis.
Get started