AI Crawlers Explained: Search, Training and User-Triggered Bots
Not every AI bot does the same job. Here are the documented categories, the agents each company publishes, and what blocking one actually costs you.
Knowledge hub
The bots behind AI search, training corpora and user-triggered fetches, and how robots.txt does and does not govern them. Under RFC 9309 those rules are requests honoured voluntarily, not access control.
Filtered to this topic. Use the row above to widen or change it.
Not every AI bot does the same job. Here are the documented categories, the agents each company publishes, and what blocking one actually costs you.
Anthropic runs three bots and states why IP blocking is the wrong opt-out. Here is each one's purpose and the documented cost of disabling it.
OpenAI runs four documented agents with independent robots.txt settings. Here is what each does, and exactly what blocking it costs you.
Perplexity runs two agents, and states that one of them generally ignores robots.txt. Here is what each does and how to configure for it.
What robots.txt can and cannot do for AI crawlers, worked policies you can copy, and how WordPress serves the file virtually.
A licensing decision, not an SEO one. What training crawlers are, what blocking actually changes, and a framework for deciding by site type.
llms.txt is a community proposal for a Markdown site map aimed at LLMs. Here is the format, its real status, and what publishing one does not do.