AI web crawlers
Bots from model labs and answer engines that fetch pages for training, retrieval, or user-initiated reading.
Name them in logs. Policy them in robots.txt one by one. “AI bot” is not a user-agent.
Bandwidth can hurt. Rate limits for scrapers vs allowlists for labs is normal ops.
Examples
- ClaudeBot fetches /docs nightly. Bytespider fetches everything twice. Different responses are allowed.
Related terms
Training bots build weights. Search/citation bots fetch for an answer now. User-triggered bots fetch because a human asked.
OpenAI runs several bots: GPTBot for training, OAI-SearchBot for ChatGPT Search, and ChatGPT-User for user-triggered fetches.
robots.txt is the fetch permission file at the site root. AI crawlers read it too. Misconfigure it and you vanish from training, search, or both.
FAQ
Are they 95% of crawl traffic? +
On some sites, AI bots dominate. Measure yours. Don’t quote someone else’s pie chart as yours.
Track this in Reddex
See Reddit threads and AI answers for your brand in one place. Start with a free analysis.
Get started