Robots.txt & AI Crawler Studio
Build clean robots.txt directives for traditional search crawlers, AI search systems and model-training crawlers while keeping crawling, indexing and AI usage policies clearly separated.
Crawler Policy Builder
Define your base crawling policy, path rules and crawler-specific access preferences.
Use URL paths, not full domain URLs.
Product token; it does not control Google Search inclusion.
Generated robots.txt
Review every directive before replacing the live file at your site’s root.
# ToolItFast robots.txt policy
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
Sitemap: https://example.com/sitemap.xml
Crawler Policy Blueprint
A compact summary of the crawl policy currently being built.
Robots.txt Policy Checks
Deterministic checks for syntax, path structure, conflicts and implementation readiness.
Generate the policy to inspect supported robots.txt directives.
Checks that crawler path rules use site-relative paths.
Looks for obvious contradictory or high-risk access choices.
Checks whether the supplied sitemap uses an absolute URL.
Separate crawling, indexing and AI usage.
Crawling is not the same as indexing
Blocking a crawler in robots.txt prevents compliant crawling, but a URL can sometimes still be discovered or referenced. Use the appropriate indexing or access controls when the goal is removal from search or true privacy.
AI search and training are different
OpenAI and Anthropic publish separate crawler identities for search-oriented access and model-development crawling. Treat those purposes independently rather than using one blanket assumption for every AI bot.
User-triggered agents can behave differently
Some services use separate agents when a human explicitly asks an AI product to retrieve a page. Those requests may have different robots.txt behavior from automated web crawling.
Publish at the site root
A robots.txt file normally belongs at the top level of the applicable host, such as example.com/robots.txt. Rules apply to the host, protocol and port where that file is served.
Verify the live file
After implementation, open the real robots.txt URL directly and verify the exact text returned by the server. CDN, plugin or hosting rules can override the file you expected to publish.
Recheck crawler documentation
AI crawler identities and provider policies can change over time. Revalidate important crawler directives against the provider’s current official documentation before major changes.
Robots.txt & AI Crawler Studio FAQ
What does a robots.txt file control?
A robots.txt file publishes crawl-access rules for compliant automated agents. It can allow or disallow paths for specific user agents and can advertise sitemap locations.
Does blocking Googlebot remove a page from Google?
Not necessarily. robots.txt primarily controls crawling. If your objective is preventing a page from appearing in search results, use an appropriate noindex or access-control method rather than relying on a crawl block alone.
Are GPTBot and OAI-SearchBot the same crawler?
No. OpenAI documents them for different purposes. OAI-SearchBot is associated with ChatGPT search discovery, while GPTBot is the crawler used for content that may contribute to model training.
Does blocking Google-Extended block Google Search?
No. Google documents Google-Extended as a separate product token for Gemini-related training and grounding controls. It does not control inclusion in Google Search and is not a Google Search ranking signal.
Can robots.txt protect private or confidential content?
No. The robots.txt file itself is public, and crawl directives are not authentication. Sensitive content should be protected with real access controls such as authentication or server-side authorization.
Does this tool send my crawler policy to an AI service?
No AI request is required for this tool. Policy generation, formatting and deterministic checks are designed to run directly in the browser.