01A public archive, and a partial one
The Internet Archive, a non-profit library in San Francisco, has been crawling the public web since 1996 and opened the Wayback Machine to everybody in 2001. Its collection is made by its own crawlers, by partner crawls, and — since 2013 — by anyone who presses Save Page Now. That is why it is so uneven: a famous home page is captured many times a day, a small business's a handful of times a decade, and plenty of names never at all.
So an empty year means nobody saved the page that year. It does not mean the site was down, and a name with no captures is not proof that it was never used. Site owners can also ask for their pages to be excluded, pages behind a login are never captured, and a site built entirely in script is sometimes captured as an empty shell. Every count on this page is a count of what the archive holds — never of what existed.
02What a capture is, and what a change is
Behind the Wayback Machine sits an index — the CDX server — with one line per capture: the moment it was taken, to the second and in UTC; the HTTP status the crawler received; the kind of content; and a digest, a fingerprint of the exact bytes that came back. This page asks that index about the name's home page.
A change point is a capture whose digest differs from the one before it. That is a change in bytes, not necessarily in meaning: a redesign is a change, and so is a new date in a footer, a rotated advert or a fresh security token. Many changes in a year says the page was busy, or dynamic; few says it stood still — or that few people saved it.
03Captures from before the registration
The creation date a registry publishes is the date of the current registration, not the first one ever. When a name lapses and is registered again — by somebody new, or by the old owner after letting it go — the date starts over. So when the archive holds captures older than that date, the name was in use before this registration: a previous owner, or a lapse. The archive cannot say which, and neither can this page.
It is worth knowing before you buy a name, or when you are trying to understand one. A name with a past can arrive with backlinks and search reputation, good or bad, with old email still being sent to it, and occasionally with a trademark dispute attached. The change points either side of the gap are usually the quickest way to see what it was: a business, a blog, a parking page with a price on it.
Registration histories — who held a name, and when — are not public. Commercial databases sell reconstructions of them; this page does not use them, and does not guess.
04What the crawler was served
Each year's bar is split by the class of answer the archive recorded. Read together they tell a story: a long run of pages, then redirects, is often a name folded into another brand; refusals and missing pages are often a site winding down, or blocking crawlers.
| Status | What it means |
|---|---|
200 | A page was served — this is the capture you can read. |
301 302 307 308 | A redirect: to https, to www, or to another name altogether. The snapshot shows where it went. |
403 404 410 | Refused or missing. A server that turned the crawler away, or a page that was not there. |
5xx | The server failed at the moment of the capture. |
- | No status recorded — usually a “revisit” record, which points back to an earlier capture with identical content. |
05How this page works
- The name you give is reduced to its registrable domain —
https://www.bbc.co.uk/newsbecomesbbc.co.uk— using the Public Suffix List. - This site's server asks the archive's CDX index for every capture of that home page, collapsed where the content did not change, and counts them by year and by status.
- At the same moment the WHOIS engine asks the registry for the name's current record — the creation date, the registrar and the expiry.
- Both are drawn on one time scale. Captures older than the creation date are drawn hatched and counted, and said so in words.
The archive is slow — eight to sixteen seconds is normal — so each answer is cached for a day, each address is limited to six readings a minute, and the archive itself is asked no more than ten times a minute from here. No archived page is embedded or shown as an image: every snapshot is a link to web.archive.org, where the archive shows it in its own frame, on its own terms. No affiliate links, no account, and nothing you look up is kept beyond the cache.