The State of Web Crawling: Inside a 66-Million-Host Database

Our crawler statistics dashboard takes a snapshot of the host database every hour. This article is the state of the crawl as of the last snapshot at writing time: 04.09.2026 15:00:00 (about 1 hour ago). It is a look inside a crawler database that has been accumulating hosts for years — what the web looks like when you actually try to count it.
The Numbers at a Glance
| Metric | Value |
|---|---|
| Hosts | 66,346,142 |
| Indexed | 27,398,114 (209,415 today) |
| Audited DNS | 16,746,541 (194,404 today) |
| Audited TLS | 9,823,991 (374,587 today) |
| Audited robots.txt | 8,705,582 (265,863 today) |
| Crawls | 6,213 (205 today) |
| Excluded | 50,738,878 |
| Data size | 43.3 GB |
About two out of five discovered hosts are indexed (27.4M of 66.3M, 41%). The rest is either still in the processing pipeline or excluded from indexing outright — and the excluded group is by far the largest, which says a lot about the real shape of the web.
Where the Hosts Live
| Region | Hosts |
|---|---|
| global | 54,916,785 |
| eu | 6,289,303 |
| as | 3,835,425 |
| af | 407,328 |
| sa | 353,268 |
| na | 284,452 |
| oc | 264,660 |
Most hosts carry no regional signal at all. Of the 11.4M that do, more than half are European (6.3M), with Asia second at 3.8M. Africa, South America, North America and Oceania each hold only a few hundred thousand identified hosts.
Suffix Distribution
| Suffix | Hosts |
|---|---|
| .com | 42,606,157 |
| .blogspot.com | 3,721,075 |
| .net | 3,309,924 |
| .cn | 1,780,703 |
| .nl | 1,470,253 |
| .icu | 1,433,179 |
| .tips | 1,048,001 |
| .org | 1,028,926 |
| .cz | 663,659 |
| .de | 601,880 |
| .ir | 536,479 |
| .co.uk | 404,974 |
.com alone covers 64% of all hosts. The runner-up is not a TLD at all but blogspot.com — millions of auto-generated Blogger subdomains. The cheap-TLD pattern (.icu, .tips, .top) is a classic bulk-registration footprint, and .cn is the largest country code in the table.
What the Host Graph Is Made Of
| Kind | Hosts |
|---|---|
| subdomain | 58,992,458 |
| apex | 7,325,905 |
| ipv4 | 30,108 |
| local | 3,761 |
| onion | 228 |
| ipv6 | 2 |
| Scheme | Hosts |
|---|---|
| http | 51,630,014 |
| https | 14,723,100 |
Almost 89% of discovered hosts are subdomains, not registered domains — the graph grows mostly by adding labels, not by buying names. The scheme split is rougher than most people expect: only 22% of discovered hosts are known under https. This is the raw discovery graph, not the audited web, and it is still dominated by plain http links.
Host Health
| Health | Hosts |
|---|---|
| excluded | 50,739,899 |
| reachable | 5,368,324 |
| alive | 3,112,402 |
| dead | 2,995,446 |
| discovered | 1,489,942 |
| unavailable | 1,372,536 |
| parked | 622,597 |
| degraded | 436,393 |
| misconfigured | 204,244 |
| invalid_url | 13,770 |
| unprocessable | 1,633 |
Of the 15.6M hosts that are not excluded, 8.5M are alive or reachable — and 3.0M are confirmed dead, nearly matching the alive group one to three. Another 1.5M sit in discovered, waiting for their first audit. The web churns constantly, and the health table is where that churn becomes visible.
The Excluded Half of the Web
| Reason | Hosts |
|---|---|
| subdomain_trap | 48,755,256 |
| crawl_canonical_redirect | 748,767 |
| crawl_robots_txt_blocked | 531,616 |
| http_3xx_redirect | 401,111 |
| crawl_noindex | 197,307 |
| blocked | 78,943 |
| crawl_duplicate_content | 25,808 |
| no_follow | 70 |
The 50.7M excluded hosts are not a rounding error — they are the majority of the database. The single dominant reason is subdomain_trap: hosts that exist only as auto-generated subdomain ladders with no real content. In our data this mass skews heavily toward Chinese-hosted sites and cheap TLDs — .cn alone holds 1.78M discovered hosts, and .icu, .top or .com.cn add hundreds of thousands more. The remaining exclusions are honest signals: canonical redirects, robots.txt blocks, HTTP redirects and noindex tags from site owners who do not want to be crawled.
Name Server Providers
| Provider | Hosts |
|---|---|
| other | 4,913,304 |
| cloudflare | 2,312,051 |
| godaddy | 440,938 |
| aws | 373,727 |
| self-hosted | 332,451 |
| alibaba | 227,748 |
| ovh | 107,848 |
| 95,521 | |
| ionos | 86,668 |
| tencent | 82,777 |
This table counts only hosts with an identified provider — 56.7M hosts carry no NS data at all (missing or none), which mirrors the exclusion mass. Among the identified ones, Cloudflare is the largest named provider by a wide margin, followed by GoDaddy and AWS. Alibaba and Tencent together serve ~310k hosts, another echo of the Chinese share of the graph.
What Fails: Top 10 Audit Error Codes
| Code | Hosts |
|---|---|
| dns_nxdomain | 4,067,545 |
| dns_noanswer | 682,526 |
| tcp_timeout | 248,623 |
| dns_servfail | 132,360 |
| tcp_unreachable | 124,130 |
| dns_timeout | 110,899 |
| http_h2_protocol_error | 82,267 |
| tls_handshake_failed | 51,561 |
| tcp_refused | 44,842 |
| dns_resolution_failed | 25,183 |
The database currently records 5,671,631 audit errors, and the top 10 codes cover 98% of them. The story is unambiguous: 72% of all failures are dns_nxdomain — domains that simply do not exist anymore — and roughly nine in ten errors are DNS-related in some form. Transport and TLS problems are a distant second. Most of the failure mass is not broken infrastructure; it is the accumulated debris of the web.
Robots Verdicts
| Verdict | Hosts |
|---|---|
| allowed | 5,658,486 |
| absent | 2,182,424 |
| blocked | 544,105 |
| unreachable | 320,567 |
Of the 8.7M hosts with a completed robots.txt audit, 65% explicitly allow crawling and another 25% ship no robots.txt at all, which defaults to allowed. Only 544k hosts actively block crawlers — a reminder that the vast majority of the web is, by its own declaration, open for indexing.
The Bigger Picture
The hourly trends behind the snapshot show steady, linear growth: roughly 11,000 new hosts and 13,000 newly indexed pages every hour, day after day. Nothing explodes, nothing collapses — a mature crawl in equilibrium between discovery, auditing and exclusion.
The full dashboard, including the hourly trend table, is public at crawler.pceuropa.net/stats and updates every hour. If you want the economics behind this kind of operation, see why we recommend Preemptible/Spot VMs for web crawling.