Test-time compute
Spending more compute at inference to think longer. Reasoning models. Latency and cost go up; answers may get more careful.
Your page still has to be retrieved. Extra thinking won’t invent a source that wasn’t fetched.
Users may not wait. Some products hide the thinking.
Examples
- A slow, careful answer still cites a stale Reddit thread because retrieval proposed it first.
Related terms
Inference is running the trained model to produce an answer. Every ChatGPT reply is inference. It costs money, which is why vendors cache and skip retrieval.
Models that spend more inference compute to think longer (o-series, R1, etc.). They can pick sources more carefully — or overthink a Reddit anecdote.
A second-stage model that reorders retrieved chunks before the LLM writes. You need to get into the first-stage set or rerank never sees you.
FAQ
GEO effect? +
Possibly fewer sloppy recommendations. Not a reason to delay publishing facts.
Track this in Reddex
See Reddit threads and AI answers for your brand in one place. Start with a free analysis.
Get started