Citation crawlers vs training crawlers
Training bots build weights. Search/citation bots fetch for an answer now. User-triggered bots fetch because a human asked.
Policy them separately in robots.txt. Mixing them is the usual foot-gun.
User-triggered fetches can still hammer you. Rate-limit without blocking the lab’s official ranges.
Examples
- You Disallow GPTBot (training) and Allow OAI-SearchBot. Search citations still possible. Memory will drift.
Related terms
Bots from model labs and answer engines that fetch pages for training, retrieval, or user-initiated reading.
OpenAI runs several bots: GPTBot for training, OAI-SearchBot for ChatGPT Search, and ChatGPT-User for user-triggered fetches.
robots.txt is the fetch permission file at the site root. AI crawlers read it too. Misconfigure it and you vanish from training, search, or both.
FAQ
Which one gets me cited this week? +
Citation/search crawlers and live browsing. Training is a slower, blurrier effect.
Track this in Reddex
See Reddit threads and AI answers for your brand in one place. Start with a free analysis.
Get started