Guide
Common Crawl and training presence
Common Crawl is a public web archive many ML datasets draw from. Presence is a weak but real signal that your domain has been seen at internet scale.
What the signal is
BatSignal checks whether your domain appears in Common Crawl-derived indexes as a coarse training-presence signal. It answers: “Has a major open crawl archive observed this site?”—not “Is the brand in every model’s weights?”
What it is not
- Not proof of inclusion in any specific closed model
- Not a substitute for live ChatGPT web-search presence
- Not something you “optimize” with keyword tricks
How to treat it operationally
- Keep public pages crawlable if you want archive inclusion
- Decide CCBot policy alongside other AI crawler rules
- Invest most effort in live answer evidence (ChatGPT citations)
- Read the full weighting in the methodology
FAQ
If I am in Common Crawl, will ChatGPT recommend me?
Not necessarily. Training presence is neither required nor sufficient for live answer-engine recommendations. BatSignal weights it at only 5% of the composite score.
Can I opt out of Common Crawl?
Common Crawl respects robots.txt for CCBot. Blocking CCBot may reduce archive inclusion; it can also reduce training-data presence. Choose deliberately.