This is an old revision of the document!
Table of Contents
Website selection
There is no register of the web, so every measurement substitutes a list for the population it wants to talk about. This page is about which list: what each ranking actually measures, what it is biased towards, and which ones the field still uses. Sampling is about how you draw from a list you have already chosen — top-n versus stratified, sample size, unit, versioning. Read them together. A reader can pick CrUX, know its biases, and still produce an unreproducible top-1,000 crawl with no list id.
This page overlaps with Website classification only in that popularity is one label people attach to sites. Classification is about topic, industry and company data. Rankings are the sampling frame.
The short version, from 1,153 papers in our corpus of seven security and privacy venues that drew a population of websites, domains or web pages (see Use in Publications):
Tranco is what 2025 papers use; CrUX is what Ruth et al. found most accurate in 2022. Of 133 papers published in 2025 that sampled the web, 75 (56.4%) named Tranco and 12 (9.0%) named CrUX. Ruth et al. measured Alexa, Majestic, Umbrella, Tranco and CrUX against Cloudflare's server-side HTTP request logs in February 2022 and found CrUX the most accurate across all metrics [1Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)]. That comparison has not been re-run on the five-provider Tranco list you download today. Alexa.com was retired on 1 May 2022 (APIs 15 December 2022); 66 papers from 2023–2026 still name it, 34 of them with no version at all. Do not write “the Tranco top 1M” without a list id: the default provider set changed on 1 August 2023. Pin the id with Tranco's pin_tranco.py.
Popular Website Lists
1. CrUX (Chrome User Experience Report)
CrUX provides rank-magnitude buckets (1k, 5k, 10k, …, 5M) of popular origins based on user-initiated page loads from eligible Chrome installs — usage-statistic reporting, history sync without a passphrase, and a supported platform (not Chrome on iOS).1) Docs live as of 2026-08-27.
- Advantages: Closest to Cloudflare's HTTP request logs in February 2022, according to [1Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)]; derived from eligible Chrome page loads; includes country- and device-specific slices; CC BY 4.0; a default Tranco input since 1 August 2023.
- Limitations: Rankings are aggregated into buckets — there is no rank 5,000; opt-in Chrome, not Chrome's full install base; monthly, not daily. Galloway et al. did not evaluate it [2Galloway, Tillson; Karakolios, Kleanthis; Ma, Zane; Perdisci, Roberto; Keromytis, Angelos D.; Antonakakis, Manos (2024): "Practical Attacks Against DNS Reputation Systems", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)].
- More details, API: CrUX
2. Tranco
Tranco is a research-oriented, hardened top-sites ranking designed to reduce churn and manipulation. It was introduced in NDSS 2019 [3Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)], when it combined Alexa, Cisco Umbrella, Majestic and Quantcast; its inputs have changed since. The daily list of 2026-08-26 (id 46W9X) is built from CrUX, Farsight, Majestic, Cloudflare Radar and Cisco Umbrella, aggregated over 30 days with the Dowdall rule — Alexa is no longer among them.2)
- Advantages: Permanent list ids; research-oriented generator; Dowdall over a month is more stable than a single daily source; configurable (with a Farsight caveat, below).
- Limitations: An aggregate, not a ground truth — Ruth et al. found it less accurate than CrUX against HTTP logs [1Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)]; dependent on its inputs, which have changed three times since the 2019 configuration (Quantcast dropped 2020-04, Farsight joined 2022-05-01, Alexa replaced by CrUX and Radar on 2023-08-01); Radar's CC BY-NC 4.0 rides along on the default list; a 2019 id and a 2026 id are different instruments.
- More details, API: Tranco
3. Cloudflare Radar
Cloudflare Radar Domain Rankings ranks pay-level domains from DNS queries to 1.1.1.1, not from HTTP hits on Cloudflare-operated websites. An ordered top 100 (daily, global and per country) plus unordered buckets up to 1M (weekly). Details on Cloudflare Radar.
- Advantages: Free API (token with Radar Read); per-country top 100; includes infrastructure names a page-load ranking drops; a default Tranco input since 1 August 2023.
- Limitations: Unordered below rank 100 — there is no rank 5,000; DNS-based (see Limitations of DNS-Based Lists); CC BY-NC 4.0; no permanent id; manipulable with a cheap VPN [2Galloway, Tillson; Karakolios, Kleanthis; Ma, Zane; Perdisci, Roberto; Keromytis, Angelos D.; Antonakakis, Manos (2024): "Practical Attacks Against DNS Reputation Systems", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)].
- More details, API: Cloudflare Radar
4. Cisco Umbrella
Cisco Umbrella ranks names by DNS query traffic to OpenDNS resolvers. The daily top-1M zip was still published at the S3 URL above on 2026-08-27. The provider's own example file begins 1,com / 2,net / 3,google.com — TLDs and infrastructure names sit in the head.
- Advantages: Captures non-browser traffic; daily snapshot with dated historical URLs; still a Tranco default input.
- Limitations: DNS-based (see Limitations of DNS-Based Lists); biased to organisations that point resolvers at OpenDNS; typos and
*.ec2.internalsurvive; the published example is not a list of websites.
5. Majestic Million
Majestic ranks domains by backlinks. The Million page still advertised free search and download on 2026-08-27 (downloads.majestic.com/majestic_million.csv). CC BY 3.0, per Tranco.
- Advantages: Independent of user-traffic panels; useful when the question is about link structure; still a Tranco default input.
- Limitations: Does not measure visits. Link-graph rankings are cheap to manipulate relative to a toolbar or panel list [3Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)]. That paper did not evaluate CrUX.
6. Other
SimilarWeb: Commercial panel and crawl estimates; paywalled (some research exceptions exist). See SimilarWeb. Named by 12 of 1,153 web-sampling papers.Quantcast: Was a Tranco default input until 1 April 2020. The Measure product still exists as publisher analytics (quantcast.com/measure/redirects to/publisher/measure); it is not a public top-sites ranking you can pin. Named by 8 papers, none in 2025–2026.Alexa: Discontinued — alexa.com was retired on 1 May 2022 and the Alexa Top Sites and Web Information Service APIs on 15 December 2022.3) Rankings came from a browser toolbar / panel, i.e. opted-in visits rather than DNS [4Scheitle, Quirin; Hohlfeld, Oliver; Gamba, Julien; Jelten, Jonas; Zimmermann, Torsten; Strowes, Stephen D.; Vallina-Rodriguez, Narseo (2018): "A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists", in: Proceedings of the Internet Measurement Conference 2018, pp. 478–493. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)]. Do not start a new crawl from an Alexa snapshot unless the study is a reproduction of a pre-2022 paper, in which case name the snapshot.Farsight(DomainTools): a 1-million pay-level-domain ranking built from cache-misses in Farsight's DNSDB — “above-resolver” DNS that organisations share, so a name that stays in a recursive cache is invisible. Tranco has used it as a default input since 1 May 2022, the same day Alexa.com was retired.4) It is not a public CSV. You get it only as part of Tranco's daily list: the custom-list API'sprovidersenum iscrux | majestic | radar | umbrella | alexa | quantcast— Farsight, which is in the daily list, is not in that enum. See Tranco. Advantages: sees cache-miss DNS from participating organisations' resolvers, including names that only infrastructure looks up; Tranco's only exclusive input. Limitations: cache-miss bias (popular names are undercounted if they stay cached); biased to organisations that share with DomainTools; DNS-based; cannot be pinned independently of a Tranco daily id. Zero papers in this corpus name the Farsight ranking as asourceList. The 13 papers that name Farsight or DNSDB used the passive-DNS dataset, which is a different instrument.SecRank: Voting-based ranking from Chinese DNS, introduced in USENIX Security 2022 [5Xie, Qinge; Tang, Shujun; Zheng, Xiaofeng; Lin, Qingran; Liu, Baojun; Duan, Haixin; Li, Frank (2022): "Building an Open, Robust, and Stable Voting-Based Domain Top List", in: 31st USENIX Security Symposium (USENIX Security 22), pp. 625-642.]. Site live on 2026-08-27. Named by 4 papers. Use it when the audience is Chinese resolver traffic; it is not a drop-in Alexa replacement for a global crawl.
Best Practices
- Pick the instrument that matches the claim. Page-load questions want CrUX (or Tranco, knowing CrUX is one input). DNS / infrastructure questions can use Umbrella, Radar or SecRank, and must say so. Backlink questions can use Majestic. “The web” is not a use-case any of these lists enumerate.
- Prefer a month of ranks, or buckets, over a single daily total order unless you need one. [3Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)] is the stability argument for aggregating over 30 days; CrUX publishes magnitudes rather than a total order, which is a different kind of stability and not one this corpus re-measured.
- Cite a list identity, not a list name. A Tranco id, a CrUX month, an Umbrella dated zip, a hash of the file you archived. 52.0% of papers that named a ranking vendor in this corpus stated no version on that tuple (397 of 763). Sampling is the page for that gap.
- Do not treat Tranco as CrUX. Tranco's default list mixes a page-load ranking with three DNS rankings and a backlink ranking. Ruth et al. found CrUX closest to HTTP logs; that comparison has not been re-run on the five-provider list you download today [1Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)].
- Publish the file. A GitHub URL is not an archival artefact — see Artifacts.
- If you can, report on more than one list. 55 of 763 ranking-users (7.2%) named two or more vendors. Ruth et al. showed CrUX and Tranco are different instruments against HTTP logs [1Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)]; they did not evaluate reporting both.
Detailed Analysis
Limitations of DNS-Based Lists
DNS-based lists — Cisco Umbrella, Farsight, SecRank, and partially Cloudflare Radar and Tranco — are built from resolver observations rather than page loads. Umbrella ranks by query traffic; Radar estimates user-population size from 1.1.1.1 (see Cloudflare Radar); Farsight ranks pay-level domains by DNSDB cache-misses. The input is DNS, which introduces several limitations:
DNS lists capture queries from devices and applications that resolve names without a user looking at a page. That over-represents software updates, network configuration and telemetry:
windowsupdate.microsoft.comis queried for updates and is not a typical user-visited website.- Internal or infrastructure names, such as
ec2.internal, appear on Umbrella; Umbrella's own example file ranks the TLDscomandnetabovegoogle.com.
Galloway et al. generated names, queried them through a $10/month VPN, and “consistently achieved a ranking in the top 100,000” on Radar; Tranco then placed the same names at rank 1 million within 10 days and inside 500k within 14 days [2Galloway, Tillson; Karakolios, Kleanthis; Ma, Zane; Perdisci, Roberto; Keromytis, Angelos D.; Antonakakis, Manos (2024): "Practical Attacks Against DNS Reputation Systems", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]. On today's Radar API that top 100,000 is an unordered bucket, not a total order — see Cloudflare Radar. That is a finding about Radar (and Tranco via that input), not about every DNS ranking.
Temporal Stability and Manipulations
Alexa, when it was live, ranked from a toolbar panel. Two problems, both measured:
- Manipulations: ranks could be altered with as little as a single HTTP request [3Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)]. There was a real-world incentive to do this.
- Stability: a small panel, a large web, and a daily publish produced high churn.
Tranco was designed to reduce both by aggregating sources over 30 days. CrUX publishes buckets rather than a total order, so a name that moves inside a magnitude stays in the same published bucket; this page does not re-measure that as a stability rate. CrUX is theoretically susceptible to manipulation; it is collected from eligible Chrome installs (not Chrome's full install base). [3Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)] costed attacking Alexa as cheap; it did not evaluate CrUX. Galloway et al. also did not evaluate CrUX [2Galloway, Tillson; Karakolios, Kleanthis; Ma, Zane; Perdisci, Roberto; Keromytis, Angelos D.; Antonakakis, Manos (2024): "Practical Attacks Against DNS Reputation Systems", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)].
Possible manipulations of lists and their estimated cost according to [3Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)].
Representativeness
- CrUX ranks origins by completed page loads, which is why Ruth et al. found it closest to Cloudflare HTTP logs. Majestic ranks by backlinks. Umbrella, Radar and Farsight rank by DNS. Those are three different populations. A result on one does not transfer to the others, and a Tranco aggregate does not make them the same population.
- Of the 1,790 Alexa top-10K domains Ruth et al. could measure against Cloudflare, 70% sat in a lower rank-magnitude bucket and 27.2% were two or more orders of magnitude lower [1Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)]. “Inaccurate” is not a small permutation of the head; it is a different web.
- Scheitle et al. compared list domains against the general
com/net/orgpopulation: HTTP/2 adoption 26.6% on the Alexa Top 1M against 7.84% overall [4Scheitle, Quirin; Hohlfeld, Oliver; Gamba, Julien; Jelten, Jonas; Zimmermann, Torsten; Strowes, Stephen D.; Vallina-Rodriguez, Narseo (2018): "A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists", in: Proceedings of the Internet Measurement Conference 2018, pp. 478–493. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)]. On that indicator the head was a technological outlier. The same paper's IPv6 and CDN gaps went the same way; it is not a theorem that every later adoption measure will.
Use in Publications
The figures below come from a structured extraction over 5,859 full-text papers from CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P, 2010–2026. The population for this page is the 1,153 papers that drew at least one study population whose unit is websites, domains or web pages — 19.7% of the corpus, the same population as Sampling. Only those web-unit tuples are counted. Free-text sourceList is folded into vendor families by scripts/rank_fold.mjs before counting (multi-label: a string naming Tranco and Radar counts for both). Sentinel values are silence. 2025 and 2026 are provisional venue-years.
This replaces the borrowed Scheitle 2018 and Xie 2024 survey figures as a description of what this corpus did. Those two papers remain the right citations for the pre-corpus surveys; they are below, dated.
Counting exact strings undercounts Alexa by 89%
The work item that commissioned this section pointed at exact-string counts over all 5,712 papers that sampled anything: “custom seed list 377, Google Play 183, Tranco 119, Alexa 117” on the previous corpus. On this run the same method produces:
Exact sourceList string | Papers of 5,712 sampled |
|---|---|
| custom seed list | 516 |
| Tranco | 180 |
| Google Play | 143 |
| Prolific | 122 |
| Alexa | 51 |
Google Play is not a website ranking. It is 281 papers after folding /google play/i over all sampled papers, and 0 over the 1,153 that sampled the web. Publishing that top-four as “which lists researchers use” would have ranked an app store above Alexa. Restrict to web units and fold spellings:
| Vendor | Papers of 1,153 | Share | Exact name only | Spellings folded | Exact-string undercount |
|---|---|---|---|---|---|
| Alexa | 463 | 40.2% | 50 | 355 | 89.2% |
| Tranco | 262 | 22.7% | 178 | 91 | 32.1% |
| CrUX | 36 | 3.1% | 9 | 23 | 75.0% |
| Majestic | 26 | 2.3% | 4 | 17 | 84.6% |
| Cisco Umbrella | 24 | 2.1% | 2 | 23 | 91.7% |
| SimilarWeb | 12 | 1.0% | 9 | 5 | 25.0% |
| Quantcast | 8 | 0.7% | 4 | 5 | 50.0% |
| Cloudflare Radar | 4 | 0.3% | 3 | 4 | 25.0% |
| SecRank | 4 | 0.3% | 2 | 4 | 50.0% |
763 papers (66.2%) name at least one of these vendors. Family counts do not sum to 763: 55 papers (7.2% of vendor-users) name two or more. Farsight-as-ranking: 0 papers. Radar as a named source: 4, all in 2025–2026 — it is a Tranco input, not a list people cite. The full fold, including commercial also-rans (Semrush 3, BuiltWith 2, Open PageRank 2, Ahrefs / Moz / Netcraft / Trexa 1 each) and the unnamed-top-N residue (4 strings), is on the provenance page.
The list changed under the field
Papers naming each list, as a share of the 1,153 in that year:
| Year | Papers | Alexa | Tranco | CrUX |
|---|---|---|---|---|
| 2021 | 87 | 47.1% | 31.0% | 0.0% |
| 2022 | 106 | 37.7% | 26.4% | 1.9% |
| 2023 | 122 | 24.6% | 38.5% | 4.1% |
| 2024 | 122 | 17.2% | 45.9% | 9.8% |
| 2025 (provisional) | 133 | 8.3% | 56.4% | 9.0% |
| 2026 (provisional) | 57 | 7.0% | 38.6% | 7.0% |
Tranco overtook Alexa in 2023, the first full year after Alexa.com died. The lag is compatible with a submission cycle; we did not read why each of those papers still named Alexa (cached CSV, inherited list, or a crawl that started before May 2022 are all possible). CrUX peaked at 9.8% in 2024 and did not keep climbing. Read 2026 as incomplete: CCS and IMC 2026 have not been held.
Of the 66 papers from 2023–2026 that still name Alexa, 32 (48.5%) state a version or date and 34 (51.5%) do not. The undated 34 cannot be pinned to a snapshot from the paper. A mirror, a cached CSV, or a list inherited from earlier work would all look like this; this corpus does not distinguish them.
Umbrella, Majestic, Radar, SimilarWeb and SecRank stay in the low single digits in every year they appear. They are specialised instruments and Tranco inputs, not a second mainstream ranking.
What the pre-corpus surveys said
Scheitle et al. [4Scheitle, Quirin; Hohlfeld, Oliver; Gamba, Julien; Jelten, Jonas; Zimmermann, Torsten; Strowes, Stephen D.; Vallina-Rodriguez, Narseo (2018): "A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists", in: Proceedings of the Internet Measurement Conference 2018, pp. 478–493. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] and Xie et al. [6Xie, Qinge; Li, Frank (2024): "Crawling to the Top: An Empirical Evaluation of Top List Use", in: International Conference on Passive and Active Network Measurement, pp. 277-306.] surveyed ranking-list usage in web-measurement papers by hand. They are why this page had figures before this corpus existed. Scheitle ends in 2016-era crawls; Xie's survey of crawling papers ends around 2022. Neither includes 2023–2026, which is the window in which Tranco became the default and Alexa became residue. Alexa was not “discontinued in 2023” — that sentence was on an earlier version of this page and confused Amazon's 2022 retirement with Tranco's 2023 provider swap.
Popularity of website lists in web measurement publications according to [4Scheitle, Quirin; Hohlfeld, Oliver; Gamba, Julien; Jelten, Jonas; Zimmermann, Torsten; Strowes, Stephen D.; Vallina-Rodriguez, Narseo (2018): "A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists", in: Proceedings of the Internet Measurement Conference 2018, pp. 478–493. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)].
Popularity of website lists in web measurement publications according to [6Xie, Qinge; Li, Frank (2024): "Crawling to the Top: An Empirical Evaluation of Top List Use", in: International Conference on Passive and Active Network Measurement, pp. 277-306.].
Methodology and limitations of these figures
- How they were produced. One structured record per paper; each
population[]tuple carries a verbatim evidence quote. The population is the extraction'sunitenum (stable). The free-textsourceListis not, and is folded before counting. - Folding.
scripts/rank_fold.mjsmaps a string onto every vendor it names. Alexa undercounts by 89.2% if you count the exact string “Alexa”. Farsight/DNSDB strings are a hand map of 12 spellings classified as a passive-DNS dataset, not a ranking; the report fails if a new spelling appears. Residue of unnamed “top N websites” strings: 4. - This is not the 5,712-paper
sampledpopulation. That denominator is 97.5% of a corpus that is mostly not web measurement. Google Play, Prolific and MNIST live in it. - Sampling's 764 “popularity-ranking” papers are a frame-kind fold, not a vendor fold; the vendor union here is 763. The one-paper gap is an unnamed top-N string that names no vendor. Sampling reports the frame kind; this page reports the vendor.
- Posters. 23 of 1,153 WEB papers are posters; 10 of 763 vendor-users. Trimmed shares (posters and ≤4-page records removed, 1,113 papers) move Alexa 40.2% → 41.0% and Tranco 22.7% → 22.8%. The headline is not a silence rate.
- Venue coverage. Seven venues only. 2025 and 2026 are incomplete by construction.
- Every query, the report script and its unedited output are on website_selection; corpus-level caveats are on corpus.
What to report
A methods sentence a reader can act on names, in this order:
- Which list, with its identity — Tranco id, CrUX month, Umbrella dated zip, Majestic download date, Radar snapshot date. Not “the top 1M”.
- What that list measures — page loads, DNS queries, cache-misses, backlinks, or an aggregate of those.
- The draw — top-n, stratified, random, … — which is Sampling, not this page.
- The file — archived, hashed, and deposited. See Artifacts.
python3 pin_tranco.py on Tranco prints the live configuration and a cite line. Use it.
Related pages
- Sampling — how to draw, sample size, unit, versioning. Assumes you have chosen a list.
- Longitudinal — why a 2019 Tranco id and a 2026 Tranco id are not commensurable.
- Tranco — ids, API, the Farsight-is-default-list-only trap,
pin_tranco.py. - Cloudflare Radar — buckets are not ranks; CC BY-NC 4.0; Galloway.
- CrUX — magnitude buckets, BigQuery, country/device slices.
- SimilarWeb — the commercial panel.
- Website classification — topic labels, not popularity ranks. Alexa died as a categoriser on the same day it died as a ranking.
- Artifacts — depositing the pinned list.
- corpus — how the 5,859-paper extraction was built.
- website_selection — every query on this page, the fold, the residue, the review log.
References
- [1]
- Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
- [2]
- Galloway, Tillson; Karakolios, Kleanthis; Ma, Zane; Perdisci, Roberto; Keromytis, Angelos D.; Antonakakis, Manos (2024): "Practical Attacks Against DNS Reputation Systems", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
- [3]
- Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)
- [4]
- Scheitle, Quirin; Hohlfeld, Oliver; Gamba, Julien; Jelten, Jonas; Zimmermann, Torsten; Strowes, Stephen D.; Vallina-Rodriguez, Narseo (2018): "A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists", in: Proceedings of the Internet Measurement Conference 2018, pp. 478–493. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
- [5]
- Xie, Qinge; Tang, Shujun; Zheng, Xiaofeng; Lin, Qingran; Liu, Baojun; Duan, Haixin; Li, Frank (2022): "Building an Open, Robust, and Stable Voting-Based Domain Top List", in: 31st USENIX Security Symposium (USENIX Security 22), pp. 625-642.
- [6]
- Xie, Qinge; Li, Frank (2024): "Crawling to the Top: An Empirical Evaluation of Top List Use", in: International Conference on Passive and Active Network Measurement, pp. 277-306.
https://tranco-list.eu/api/lists/date/latest returns list_id: 46W9X, providers: [crux, farsight, majestic, radar, umbrella], combinationMethod: dowdall, filterPLD: on, listPrefix: 1000000. Describe Tranco's composition from the live API rather than from the 2019 paper. See Sampling and Tranco.


