In one sentence: Common Crawl is a nonprofit that has spent years copying public web pages and giving the resulting archive away for free, which has made it one of the most widely used raw sources of web text for building AI models.
Every so often the Common Crawl project sends out a crawler, grabs an enormous sample of public pages, and saves the text. Those snapshots are then published for anyone to download at no cost. That is the unusual part. Most crawl archives of that size are private company property, and this one is not. What gets saved is the text and markup the server sent, not your design, and not anything that only appears on the page after a script runs.
Because it is free and very large, it became a cheap starting point for teams building a language model, and for academics studying the web besides.
New snapshots keep getting added, and the old ones stay where they are. So there is no single copy of your site in there. There is a run of them going back years, each one a photograph of whatever the page said the month it was taken.
Here is the piece owners miss: it is an archive, not a mirror. It holds what your page said on the day it was copied, and nothing you do to your site reaches back and changes it. Say you moved your shop across town a few years ago and changed the phone number at the same time. A snapshot taken before the move still carries the old address and the old number, sitting in a public file anyone can download.
An assistant draws on two things: what it absorbed while it was being built, and whatever it looks up at the moment somebody asks. Archives like this one feed the first kind. If an old copy of your page is what got absorbed, an assistant can state your former hours or your former address with complete confidence, because it has no way to know the page ever changed.
You cannot edit the archive, and you cannot pull old copies back out of models that already read them. What you can do is make the current version of the truth loud, plain, and consistent everywhere a machine can reach it right now: your own pages, your listings, your profile. In practice that means the same business name, the same address, the same phone number and the same list of services written in plain text on your site and on every profile you control, so there is nothing left to disagree with. Over time, current sources that agree with each other outweigh one stale copy, and getting your real facts stated plainly and repeated consistently is most of what AI search optimization actually is.
Common Crawl is not Google's index, and it is not an AI crawler run by a model company for its own assistant. It is a third party archive that many companies happen to draw from. It is also not the whole story of what a model knows, since the archive is one input among many that end up as training data.
And it is not something you opted into. Unless somebody wrote a rule for its crawler at some point, the default was to be copied, the same as it is for every other crawler that reads the open web.
You can tell its crawler not to take new copies by naming it in your robots.txt file, the same way you would with any other crawler. That stops future snapshots. It does not erase copies already published, and it cannot reach copies other people have already downloaded.
No. You remove one source, not all of them. Assistants also read live pages at the moment of the question and pull from listings and reviews, so blocking one archive usually means the picture they hold of you is thinner rather than absent.
AI search optimization · Training data · AI crawler · Crawling · Large language model · All glossary terms · Plain-English answers
Free consultation, plain-English advice. If you don't need us, we'll say so.
Book a free consultation → Or call/text directly: (407) 694-2055Tell us a little about the business and we will come back with an honest read: what we would fix first, what it costs, and whether you need us at all. Prefer to see work before you talk numbers? Get a free homepage mockup, built for your business, yours to keep either way.
Brandon reads every one of these himself. You will hear back shortly with an honest read on what we would do first, what it costs, and whether it is worth it for you.