Content Signals (robots)
Proposed robots.txt-style directives for how crawled content may be used: search vs AI input vs training.
Early and unevenly supported. Still worth watching if you have a legal/policy team.
Until adoption is real, robots.txt user-agents remain the practical switch.
Examples
- A publisher sets search=yes, train=no. Some labs honor it, some don’t, some haven’t heard of it.
Related terms
llms.txt is a Markdown file at the site root that tells AI crawlers what the site is and which pages to read first. It is not a second robots.txt with legal teeth.
robots.txt is the fetch permission file at the site root. AI crawlers read it too. Misconfigure it and you vanish from training, search, or both.
Legal/technical flags that reserve text-and-data-mining rights. EU-flavored. Partial crawler support.
FAQ
Replace robots.txt? +
No. Additive, if anything. Keep the old file correct.
Track this in Reddex
See Reddit threads and AI answers for your brand in one place. Start with a free analysis.
Get started