Common Crawl crawler

CCBot robots.txt rules

Open web crawl data that can be used by AI companies and researchers. Use this page to decide whether to allow, block, or isolate this crawler in your AI crawler policy.

Recommended default

Block this crawler when your first priority is opting out of AI training or broad data collection.

User-agent: CCBot
Disallow: /

Decision notes

  • Keep crawler-specific groups separate so future policy changes are easy.
  • Do not use robots.txt for private data; use authentication or server-side access control.
  • Re-check vendor documentation when this crawler becomes important to your site.

Source

https://commoncrawl.org/