User Tools

Site Tools


programming:crawler:webxray

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Next revision
Previous revision
programming:crawler:webxray [2026/08/17 07:54] – New page: webXray, its domain-ownership database, and a measured three-way comparison with DuckDuckGo Tracker Radar and Disconnect entities.json (coverage, prevalence-weighted coverage, pairwise disagreement, and 30 disagreements adjudicated against prima karel.kubicek.claudeprogramming:crawler:webxray [2026/08/17 08:13] (current) – BLOCKING correction: webXray did not end up source-available. Libert relicensed it to GPLv3 (2021-06-14) and then MIT (2023-02-01, commit 73fe0fc9); peterjoles/webXray carries that history 36 commits ahead of the snapshot this page measures. The licence b karel.kubicek.claude
Line 5: Line 5:
 Two things a new measurement needs to know before citing it: Two things a new measurement needs to know before citing it:
  
-  * **The tool is gone.** ''github.com/timlib/webXray'' returns HTTP 404there has never been a PyPI package, the surviving copies are dormant, and the licence forbids redistributing it. The author now runs a commercial product under the same name. +  * **The tool is gone from where every paper points, but it is not lost and it did not end up proprietary.** ''github.com/timlib/webXray'' returns HTTP 404 and there has never been a PyPI package — but Libert relicensed webXray twice after its most widely mirrored snapshotto GPLv3 in 2021 and then to **MIT in 2023**, and that history survives in a fork. The author now runs a commercial product under the same name at ''webxray.ai''
-  * **The database survives, but it is frozen at 2021 and badly stale wherever it has been checked.** Of 30 high-prevalence ownership disagreements adjudicated against primary sources on 2026-08-17, webXray names today's owner for **3** of the 25 domains it covers at all. Those 30 were selected //because// the lists disagreed on them, so that is not an error rate — but the direction is not in doubt. Tracker Radar and Disconnect are the live alternatives, and they are not equivalent to each other either.+  * **The database survives, but it is frozen at 2021 and badly stale wherever it has been checked.** Of 28 high-prevalence ownership disagreements settled against primary sources on 2026-08-17, webXray names today's owner for **1** of the 23 domains it covers at all. Those 28 were selected //because// the lists disagreed on them, so that is not an error rate — but the direction is not in doubt. Tracker Radar and Disconnect are the live alternatives, and they are not equivalent to each other either.
  
-This page is therefore two pages in one: what webXray was and why you cannot install it, and — the part you actually need — **how domain-to-company ownership resolution works now**, measured. If your question is "which crawler should I run", go to [[Programming:Crawler]]; if it is "how do I decide which company a request went to", you are in the right place.+This page is therefore two pages in one: what webXray was and why you cannot install it, and — the part you actually need — **how domain-to-company ownership resolution works now**, measured. If your question is "which crawler should I run", go to [[Programming:Crawler]]; if it is "how do I decide which company a request went to", you are in the right place, and the short answer is in [[#Choosing a resolution source now|Choosing a resolution source now]] with the pipeline around it in [[#Assembling the pipeline|Assembling the pipeline]]. Everything between those and here is the evidence for them.
  
 <WRAP important> <WRAP important>
Line 33: Line 33:
 | ''webxray.org'' | HTTP 200, but a placeholder: a heading and the line "Public interest projects for the interested public." No source link, no version, no download. | | ''webxray.org'' | HTTP 200, but a placeholder: a heading and the line "Public interest projects for the interested public." No source link, no version, no download. |
 | PyPI ''webxray'' / ''web-xray'' / ''policyxray'' | 404 each. webXray was never packaged on PyPI. | | PyPI ''webxray'' / ''web-xray'' / ''policyxray'' | 404 each. webXray was never packaged on PyPI. |
-| ''github.com/thezedwards/webXray'' | The most complete surviving copy of webXray 3.x, last commit **2021-03-04**. Its README still instructs ''git clone https://github.com/timlib/webXray.git''. | +| ''github.com/thezedwards/webXray'' | The most widely mirrored copy of webXray 3.x, last commit **2021-03-04**. Its README still instructs ''git clone https://github.com/timlib/webXray.git''. Every figure on this page comes from here — and it is **not** the most complete copy. | 
-the 19 forks of that copy | all dormant; the newest activity anywhere in the network is 2023-03-12, three small commits in a personal working copy.((''api.github.com/repos/thezedwards/webXray/forks?per_page=100'', checked 2026-08-17. Most recent by ''pushed_at'': ''peterjoles/webXray'' 2023-03-12.)) |+''github.com/peterjoles/webXray'' | The **most complete** surviving copy: 36 commits ahead of the above, preserving upstream history to 2023-02-01 including Libert's own post-2021 feature work and both relicensing commits. Last push 2023-03-12. This is the copy to take. | 
 +| the fork network | dormant; nothing pushed anywhere since 2023-03-12. ''forks_count'' reports 19 while the forks endpoint returns 20 objects.((''api.github.com/repos/thezedwards/webXray'' and ''.../forks?per_page=100'', checked 2026-08-17.)) |
 | ''github.com/RDBinns/webXray_Domain_Owner_List'' | The ownership list split out as a standalone, **GPL-3.0** repository, created and last pushed on 2018-04-05. Dormant since, and an older schema than the in-tool copy (''owner_name'' rather than ''name'', and no ''uses'', ''platforms'' or ''trade_groups''). | | ''github.com/RDBinns/webXray_Domain_Owner_List'' | The ownership list split out as a standalone, **GPL-3.0** repository, created and last pushed on 2018-04-05. Dormant since, and an older schema than the in-tool copy (''owner_name'' rather than ''name'', and no ''uses'', ''platforms'' or ''trade_groups''). |
 | ''webxray.ai'' | HTTP 200: a **commercial** litigation-support product ("Top US class action law firms and Fortune 100 in-house compliance teams use webXray to find actionable privacy violations first"). Libert's own homepage states "(Dr.) Timothy Libert is founder and CEO of webXray LLC."((''https://webxray.ai/'' and ''https://timlibert.me/'', both fetched 2026-08-17. The ''timlib'' GitHub profile lists ''company: webXray.ai''.)) | | ''webxray.ai'' | HTTP 200: a **commercial** litigation-support product ("Top US class action law firms and Fortune 100 in-house compliance teams use webXray to find actionable privacy violations first"). Libert's own homepage states "(Dr.) Timothy Libert is founder and CEO of webXray LLC."((''https://webxray.ai/'' and ''https://timlibert.me/'', both fetched 2026-08-17. The ''timlib'' GitHub profile lists ''company: webXray.ai''.)) |
  
 <WRAP important> <WRAP important>
-**webXray is not open source, and redistributing it is prohibited.** The surviving copy'''LICENSE.md'' is the **PolyForm Strict License 1.0.0**, which grants use for any noncommercial purpose — explicitly including "public research organization" and "educational institution" — but grants **no** right to distribute the software or to make "changes or new works based on the software". PolyForm'own summary is blunter: Strict "removes permission to distribute copies and make changes, leaving only permission to use for noncommercial purposes". 1.0.0 is still the only version, and note that PolyForm Strict has **no SPDX identifier** — if your artefact metadata expects one, there is none to give.((''https://polyformproject.org/licenses'' and ''.../licenses/strict/1.0.0'', checked 2026-08-17: ''strict/1.0.0'' is the only Strict version listed. SPDX's licence list carries ''PolyForm-Noncommercial-1.0.0'' and ''PolyForm-Small-Business-1.0.0'' but no Strict entry, checked against ''spdx/license-list-data'' on the same date.)) Its README says the same in plainer words: "This software is //not// open sourceit is //source available// and licensed for non-commercial use onlyYou may not distribute webXray in whole or in part or sell data generated by webXray without prior written permission."+**Check the licence of the copy you actually take: webXray'licence changed three times and the surviving copies disagree.** This is the most misleading thing about webXray'remains, and it caught this page — the first version asserted from one snapshot that webXray "is not open source and redistributing it is prohibited". That is true of that snapshot and false of the project's final state.
  
-The practical consequence is not academicUpstream is a 404so the only remaining route to the code is a redistribution the licence forbidsTreat webXray as **unobtainable** and do not plan study around it. Even if you obtain itit no longer installs cleanly: ''requirements.txt'' pins ''lxml==4.6.2'' and ''psycopg2-binary==2.8.6'', and neither has a PyPI wheel beyond CPython 3.9 — which has been end-of-life since October 2025 — so ''pip install -r requirements.txt'' on a current interpreter falls back to source builds that need matching ''libxml2''/''libxslt'' and PostgreSQL headers.((Checked on PyPI, 2026-08-17: the newest wheels for ''lxml'' 4.6.2 and ''psycopg2-binary'' 2.8.6 are ''cp39''. ''websocket-client'' 0.57.0 and ''textstat'' 0.7.0 ship universal/py3 wheels and are fine.)) The 2018 standalone ownership list at ''RDBinns/webXray_Domain_Owner_List'' is a separate matter: its own README licenses it under **GPLv3**, so that snapshot can be used and redistributed — but it is the 2018 schema, not the 2021 file every figure on this page is measured from. Verify the licence of whatever file you actually download rather than assuming one licence covers the project.+^ Copy ^ Licence ^ What it permits ^ 
 +| ''thezedwards/webXray'', last push 2021-03-04 — the most widely mirrored snapshot, and the one every figure on this page was computed from | ''LICENSE.md'' is **PolyForm Strict 1.0.0**; its README says "not open source… source available… non-commercial use only" | use for any noncommercial purpose, explicitly including "public research organization" and an "educational institution" — but **no** right to distributeor to make "changes or new works based on the software". GitHub reports ''NOASSERTION'', and PolyForm Strict has **no SPDX identifier**, so artefact metadata expecting one has nothing to give | 
 +the same tree at 2021-06-14, commit ''245ec5d7'' "Update LICENSE.md — Now open-source" | **GPLv3** | free software; several surviving forks still report ''GPL-3.0''
 +the last state, commit ''73fe0fc9'' of 2023-02-01, authored by Tim Libert and visible today at ''peterjoles/webXray'' | **MIT**, "Copyright (c) 2023 Tim Libert" | anything, including redistributionIts README still says "GPLv3, open source" — the README lagged the file | 
 + 
 +So webXray ended **MIT-licensed**and the fork carrying that history is 36 commits ahead of the copy measured here.((''api.github.com/repos/peterjoles/webXray'' reports ''spdx_id: MIT'' and ''pushed_at: 2023-03-12''; ''compare/master...peterjoles:master'' reports ''ahead_by: 36''; the ''LICENSE'' commit ''73fe0fc9'' is authored by "Tim Libert". Checked 2026-08-17. Also checked: ''polyformproject.org/licenses'' lists ''strict/1.0.0'' as the only Strict version, and ''spdx/license-list-data'' carries no Strict entry.)) That does not make webXray live — upstream is still 404, no fork has been pushed since 2023-03-12, and it still will not install: ''requirements.txt'' pins ''lxml==4.6.2'' and ''psycopg2-binary==2.8.6'', neither of which has a PyPI wheel beyond CPython 3.9end-of-life since October 2025.((Checked on PyPI, 2026-08-17. ''websocket-client'' 0.57.0 and ''textstat'' 0.7.0 ship universal/py3 wheels and are fine.)) But it changes what you may lawfully do with it, and it means **every figure here is measured from the most restrictively licensed copy in existence**. If you reuse webXray, take the 2023 MIT tree and name the commit. 
 + 
 +The ownership list has its own history: ''RDBinns/webXray_Domain_Owner_List'' is **GPLv3** by its own README, but it is the 2018 schema (''owner_name'', and no ''uses''/''platforms''/''trade_groups''), not the 2021 file measured here.
 </WRAP> </WRAP>
  
Line 75: Line 83:
 ===== How it compares to Tracker Radar and Disconnect ===== ===== How it compares to Tracker Radar and Disconnect =====
  
-The three lists are not three attempts at the same artefact. They differ in size by more than an order of magnitude, in what a record means, and in what they are licensed for.+The three lists are not three attempts at the same artefact. They differ in size by more than an order of magnitude, in what a record means, and in what they are licensed for. A fourth live option, Ghostery's ''trackerdb'', is **not** measured here: its ownership data is spread across per-company ''.eno'' files with a separate ''patterns'' layer rather than a single domain→owner map, so putting it in the same table would have meant writing a parser whose choices nobody could check against the other three. That is an omission, not a judgement — if you are choosing among the live lists, this page gives you numbers for two of the three.
  
 ^ ^ webXray ''domain_owners.json'' ^ DuckDuckGo Tracker Radar ''entity_map.json'' ^ Disconnect ''entities.json'' ^ ^ ^ webXray ''domain_owners.json'' ^ DuckDuckGo Tracker Radar ''entity_map.json'' ^ Disconnect ''entities.json'' ^
Line 96: Line 104:
 Size is the wrong comparison, because a domain you never meet costs you nothing. The right one is: **of the third-party domains a crawl actually encounters, weighted by how often it encounters them, what share can this list attribute to a company?** Size is the wrong comparison, because a domain you never meet costs you nothing. The right one is: **of the third-party domains a crawl actually encounters, weighted by how often it encounters them, what share can this list attribute to a company?**
  
-The measurement below takes Tracker Radar's ''domain_summary.json'' as the universe of third-party domains and its ''prevalence'' field as the weight, folds each key to a registrable domain with the current Public Suffix List, and asks each list for an owner. This universe is Tracker Radar's own view of the web, which flatters Tracker Radar and nobody else — read its column as "the list scored on its home ground" and the other two as measured against a denominator they had no part in choosing.+The measurement below takes Tracker Radar's ''domain_summary.json'' as the universe of third-party domains and its ''prevalence'' field as the weight, folds each key to a registrable domain with the **ICANN section** of the current Public Suffix List, and asks each list for an owner. This universe is Tracker Radar's own view of the web, which flatters Tracker Radar and nobody else — read its column as "the list scored on its home ground" and the other two as measured against a denominator they had no part in choosing.
  
-^ List ^ Domains it can name an owner for ^ Share of 45,525 ^ Weighted by prevalence ^ +^ List ^ Domains it can name an owner for ^ Share of 32,369 ^ Weighted by prevalence ^ 
-| webXray | 657 1.4% | **54.5%** | +| webXray | 669 2.1% | **58.6%** | 
-| Tracker Radar | 5,539 12.2% | **79.3%** | +| Tracker Radar | 5,581 17.2% | **84.3%** | 
-| Disconnect (''properties'' ∪ ''resources'') | 2,277 5.0% | **75.8%** |+| Disconnect (''properties'' ∪ ''resources'') | 2,268 7.0% | **80.5%** |
  
-The 1.4% and the 54.5% in the same row are the whole story of this file. webXray covers almost none of the web by domain count and more than half of it by weight, because it is a **head list**: the few hundred domains it knows are the ones on every page. Sliced by rank, the tail falls off a cliff:+The 2.1% and the 58.6% in the same row are the whole story of this file. webXray covers almost none of the web by domain count and well over half of it by weight, because it is a **head list**: the few hundred domains it knows are the ones on every page. Sliced by rank, the tail falls off a cliff:
  
 ^ Slice of the universe, by prevalence ^ webXray ^ Tracker Radar ^ Disconnect ^ ^ Slice of the universe, by prevalence ^ webXray ^ Tracker Radar ^ Disconnect ^
-| top 100 domains | 70% | 95% | 91% | +| top 100 domains | 71% | 98% | 94% | 
-| top 1,000 domains | 26% | 67% | 67% | +| top 1,000 domains | 26% | 69% | 68% | 
-| top 10,000 domains | 5% | 27% | 16% |+| top 10,000 domains | 5% | 29% | 17% |
  
 Libert said this himself in 2018, and the sentence is worth having to hand when a reviewer asks about coverage: "because webxray's database of domain ownership primarily contains major ad networks rather than small clients, and policyxray only searches for identified parties, variability in the long-tail of trackers may not have an outsized effect on overall findings related to disclosure. Nonetheless, it is important to point out that the number of parties being searched for is fewer than the total number of parties present." {[libert2018_automated]} Libert said this himself in 2018, and the sentence is worth having to hand when a reviewer asks about coverage: "because webxray's database of domain ownership primarily contains major ad networks rather than small clients, and policyxray only searches for identified parties, variability in the long-tail of trackers may not have an outsized effect on overall findings related to disclosure. Nonetheless, it is important to point out that the number of parties being searched for is fewer than the total number of parties present." {[libert2018_automated]}
  
-Two traps in the denominatorboth of which will bite anyone who repeats this measurement:+Three traps hereand the first one will silently ruin any coverage figure you compute:
  
-  * **Tracker Radar'''domain_summary.json'' is keyed by //hostname// for 16,396 of its 47,836 rows**, so ''fonts.googleapis.com'', ''ajax.googleapis.com'' and ''maps.googleapis.com'' are separate rows while ''googleapis.com'' is not one at all. All three ownership lists key on the registrable domain, so comparing coverage without folding to eTLD+1 charges webXray and Disconnect for subdomains they were never meant to hold. After folding, ''fonts.googleapis.com'' remains the single most prevalent third-party name in the data (0.369) with **no** owner in any of the three lists.+  * **The Public Suffix List has two sections, and using the wrong one manufactures a coverage hole.** ''googleapis.com'' is **PRIVATE** PSL rule (6,941 ICANN rules against 3,290 private ones in the snapshot used here). Fold with the private section and ''fonts.googleapis.com'', ''ajax.googleapis.com'' and ''maps.googleapis.com'' each survive as their own "registrable domain"; look each up by exact key and every one comes back unowned — even though **all three lists name ''googleapis.com'' → Google**. ''fonts.googleapis.com'' alone carries prevalence 0.369, so this single mistake moves webXray's weighted coverage from 58.6% to 54.5% and Tracker Radar's from 84.3% to 79.3%. Fold on the **ICANN section**, or walk parent labels at lookup time, or both. The four combinations are printed side by side by the script and the spread between them is larger than the difference between the lists. 
 +  * **A large fraction of ''domain_summary.json'''s 47,836 rows are keyed by hostname, and //how// large depends on the fold.** An ICANN-section fold merges **15,651** of them; the ICANN+PRIVATE fold merges only 2,339, because ''googleapis.com'', ''s3.amazonaws.com'' and ''cloudfront.net'' are themselves private-section suffixes. (A label-count heuristic — "three or more labels" — gives 16,396 and answers neither question; do not use it.All three ownership lists key on the registrable domain, so comparing coverage without folding at all charges webXray and Disconnect for subdomains they were never meant to hold.
   * **19 rows are not hostnames at all** — 18 bracketed IPv6 literals and the literal string ''"null"'', the last with a prevalence of 0.024 and a full behaviour profile attached. Drop them explicitly and say you did; do not let a ''null'' domain become a data point.   * **19 rows are not hostnames at all** — 18 bracketed IPv6 literals and the literal string ''"null"'', the last with a prevalence of 0.024 and a full behaviour profile attached. Drop them explicitly and say you did; do not let a ''null'' domain become a data point.
 +
 +With the fold and the lookup done properly, **29,896 of the 32,369 domains (92.4%) have no owner in either webXray or Disconnect, but only 14.4% of the prevalence weight does**. The most requested domain no list can name is ''tiktokw.us'' at prevalence 0.041 — Tracker Radar attributes it to ByteDance, the other two have nothing. That is what a real coverage hole looks like: recent, mid-tail, and concentrated in domains that appeared after the list was last curated.
  
 ==== Two lists disagree: is that an error, or a different question? ==== ==== Two lists disagree: is that an error, or a different question? ====
  
-For every domain that two lists both cover, ''owner_dbs.py'' compares the owner strings after folding away legal-form suffixes (''Inc'', ''LLC'', ''GmbH'', ''S.A.S'', …) but **no** synonyms — so the "agree" columns are a lower bound on real agreement and the last column an upper bound on real disagreement.+For every domain that two lists both cover, ''owner_dbs.py'' compares the owner strings after normalising punctuation and folding away legal-form suffixes (''Inc'', ''LLC'', ''GmbH'', ''S.A.S'', …) but **no** synonyms — so the "agree" columns are a lower bound on real agreement and the last column an upper bound on real disagreement. The fold is deliberately timid: it leaves **782 of webXray's 827** owner names untouched, and of the 45 it does change only **16** lose a legal suffix — the other 29 change on punctuation alone (''AT&T'', ''JD.com'', "Here, There & Everywhere"). If you want the agreement figures to go up, a synonym table is what you would have to add, and there is no principled one.
  
 ^ Pair ^ Domains both cover ^ Same after suffix fold ^ One name contains the other ^ Neither ^ ^ Pair ^ Domains both cover ^ Same after suffix fold ^ One name contains the other ^ Neither ^
-| webXray vs Tracker Radar | 601 304 (50.6%) | 100 (16.6%) | **197 (32.8%)** | +| webXray vs Tracker Radar | 612 305 (49.8%) | 109 (17.8%) | **198 (32.4%)** | 
-| webXray vs Disconnect | 461 212 (46.0%) | 33 (7.2%) | **216 (46.9%)** | +| webXray vs Disconnect | 464 213 (45.9%) | 34 (7.3%) | **217 (46.8%)** | 
-| Tracker Radar vs Disconnect | 1,532 630 (41.1%) | 318 (20.8%) | **584 (38.1%)** |+| Tracker Radar vs Disconnect | 1,538 634 (41.2%) | 320 (20.8%) | **584 (38.0%)** |
  
 Between a third and a half of jointly covered domains get different company names. Before concluding that some list is broken, note what happens if you resolve webXray up its ownership tree first — the obvious fix, since webXray says "DoubleClick" where the others say "Google": Between a third and a half of jointly covered domains get different company names. Before concluding that some list is broken, note what happens if you resolve webXray up its ownership tree first — the obvious fix, since webXray says "DoubleClick" where the others say "Google":
  
 ^ Pair ^ Neither, using webXray's immediate owner ^ Neither, using the root of webXray's tree ^ ^ Pair ^ Neither, using webXray's immediate owner ^ Neither, using the root of webXray's tree ^
-| vs Tracker Radar | 197 (32.8%) | **236 (39.3%)** | +| vs Tracker Radar | 198 (32.4%) | **239 (39.1%)** | 
-| vs Disconnect | 216 (46.9%) | **228 (49.5%)** |+| vs Disconnect | 217 (46.8%) | **230 (49.6%)** |
  
 Resolving to the root makes agreement **worse**, and the reason is instructive. webXray's roots are holding companies and, worse, //historical// ones: ''google.com'' resolves to "Alphabet" (against "Google LLC" and "Google"), ''adnxs.com'' to "AT&T", ''yahoo.com'' to "Verizon", ''tapad.com'' to "Telenor", ''turn.com'' to "Singtel". Every one of those was true when the file was last curated and none is true now. So the disagreements decompose into **two independent axes**, and a paper has to state its position on both: Resolving to the root makes agreement **worse**, and the reason is instructive. webXray's roots are holding companies and, worse, //historical// ones: ''google.com'' resolves to "Alphabet" (against "Google LLC" and "Google"), ''adnxs.com'' to "AT&T", ''yahoo.com'' to "Verizon", ''tapad.com'' to "Telenor", ''turn.com'' to "Singtel". Every one of those was true when the file was last curated and none is true now. So the disagreements decompose into **two independent axes**, and a paper has to state its position on both:
Line 190: Line 201:
 </code> </code>
  
-Four rows, four different lessons''doubleclick.net'' is the granularity axis with nothing wrong anywhere''adnxs.com'' is the vintage axis, with webXray's whole chain superseded''facebook.net'' shows the two live lists disagreeing because one carries a five-year-old legal name; and ''fonts.googleapis.com'' is the single most requested third-party name in the data with **no owner in any list**, which is what a coverage figure looks like from the insideThe ''seen'' set in ''wx_chain'' is not defensive decoration — walk a hand-curated parent chain without a cycle guard and a single bad edge hangs your pipeline.+Four rows, four different lessons''doubleclick.net'' is the granularity axiswith nothing wrong anywhere''adnxs.com'' is the vintage axis, with webXray's whole chain superseded''facebook.net'' shows the two live lists disagreeing because one carries a five-year-old legal name. And ''fonts.googleapis.com'' comes back empty from all three — **which is a bug in this snippet, not a coverage hole**: every list names ''googleapis.com'' → Google, and the snippet looks up the exact key it was givenThat is the PSL private-section trap above, reproduced in twenty lines. Fix it by trying each parent label: 
 + 
 +<code python> 
 +def resolve(host, mapping): 
 +    labels = host.split('.'
 +    for i in range(len(labels) - 1):          # stop at two labels 
 +        candidate = '.'.join(labels[i:]) 
 +        if candidate in mapping: 
 +            return mapping[candidate], candidate 
 +    return None, None 
 +</code> 
 + 
 +with which ''fonts.googleapis.com'' resolves to Google in all three. Two further details worth copying: the ''seen'' set in ''wx_chain'' is not decoration — walk a hand-curated parent chain without a cycle guard and one bad edge hangs your pipeline; and Disconnect's ''resources'' are loaded before its ''properties'' so the ownership claim wins where the two disagree.
  
 ==== Which list is right, when they disagree? ==== ==== Which list is right, when they disagree? ====
  
 30 of the highest-prevalence disagreements were adjudicated by hand against primary sources — company newsrooms, SEC filings, or the domain's own legal documents — on 2026-08-17. Every row and source is in ''scripts/owner_adjudication.py'' and on the provenance page. 30 of the highest-prevalence disagreements were adjudicated by hand against primary sources — company newsrooms, SEC filings, or the domain's own legal documents — on 2026-08-17. Every row and source is in ''scripts/owner_adjudication.py'' and on the provenance page.
 +
 +Two of the 30 could not be settled from a primary source and are **excluded from the table**, leaving 28.
  
 ^ List ^ Names today's owner ^ Stale (a real former owner) ^ Granularity only ^ Outright error ^ No entry ^ ^ List ^ Names today's owner ^ Stale (a real former owner) ^ Granularity only ^ Outright error ^ No entry ^
-| webXray | | **17** | 4 | 1 | 5 | +| webXray | | **17** | 4 | 1 | 5 | 
-| Tracker Radar | | 12 | | 2 | 0 | +| Tracker Radar | | 12 | | 2 | 0 | 
-| Disconnect | **22** | 0 | 0 | 0 | | +| Disconnect | **21** | 0 | 0 | 0 | |
- +
-Read as a share of each list's own entries: webXray is current on **3 of 25 (12%)**, Tracker Radar on **9 of 30 (30%)**, Disconnect on **22 of 22 (100%)**.+
  
 <WRAP important> <WRAP important>
-These 30 rows were chosen **because** the lists disagreed on them, ranked by prevalence. They are not a random sample, so the table above is not an error rate for any list — a random sample would be dominated by domains all three get right. What it does establish, and what a random sample would not show as sharply, is the //shape// of the disagreement: it is concentrated in acquisitions and renamesit points overwhelmingly one wayand the list that is regenerated most conservatively is not the one that is most current.+**This is not an accuracy ranking, and it cannot be turned into one.** The 28 rows were chosen //because// the lists disagreed on them, ranked by prevalencea random sample would be dominated by domains all three get right, so no percentage taken from this table means anything about how often a list is correct in generalTwo further columns are artefacts of how the sample was drawn rather than results: Tracker Radar's zero in "No entry" is partly because the rows were ranked by Tracker Radar's own prevalence field, and Disconnect's zero in "Stale" is over the 21 rows it covers at all, not over 28. 
 + 
 +What the table //does// establish, and what a random sample would show less sharply, is the **shape** of the disagreement. It is concentrated in acquisitions and renames rather than spread across the list; it points overwhelmingly one wayand the list regenerated most conservatively is not the one that is most current. If you need an accuracy rate, you have to draw a stratified random sample and adjudicate it — nobody has, and this page is not a substitute.
 </WRAP> </WRAP>
  
Line 223: Line 248:
 | **DuckDuckGo Tracker Radar** ''entity_map.json'' / ''domain_map.json'' | current; regenerated monthly, last commit 2026-08-12; CC BY-NC-SA 4.0 | you want the **broadest** coverage, prevalence weights, or per-domain categories and fingerprinting scores in the same dataset. Expect legal-entity names, and expect renames to lag. | | **DuckDuckGo Tracker Radar** ''entity_map.json'' / ''domain_map.json'' | current; regenerated monthly, last commit 2026-08-12; CC BY-NC-SA 4.0 | you want the **broadest** coverage, prevalence weights, or per-domain categories and fingerprinting scores in the same dataset. Expect legal-entity names, and expect renames to lag. |
 | **Ghostery ''trackerdb''** / WhoTracks.me | current; ''trackerdb'' last commit 2026-08-06, **CC BY-NC-SA 4.0** (its ''package.json'' says so explicitly; GitHub reports no SPDX id, so do not trust the API field); WhoTracks.me data repo last commit 2026-08-04 ("July update"); the site now redirects to ''ghostery.com/whotracksme'' | you want an ''organizations'' + ''patterns'' model where one company can carry several independently categorised behaviours (Google Analytics separate from Google Tag Manager), or the {[karaj2018whotracksme]} longitudinal data | | **Ghostery ''trackerdb''** / WhoTracks.me | current; ''trackerdb'' last commit 2026-08-06, **CC BY-NC-SA 4.0** (its ''package.json'' says so explicitly; GitHub reports no SPDX id, so do not trust the API field); WhoTracks.me data repo last commit 2026-08-04 ("July update"); the site now redirects to ''ghostery.com/whotracksme'' | you want an ''organizations'' + ''patterns'' model where one company can carry several independently categorised behaviours (Google Analytics separate from Google Tag Manager), or the {[karaj2018whotracksme]} longitudinal data |
-| **webXray ''domain_owners.json''** | **historical**; frozen at 2021-03-04no update path, tool unobtainable | reproducing or extending a pre-2022 result that used it, or you specifically need the ''parent_id'' tree or the per-language policy URLs and will re-verify each owner you rely on |+| **webXray ''domain_owners.json''** | **historical**; the file is frozen at 2021-03-04 with no update path, and the tool has had no commit since 2023 | reproducing or extending a pre-2022 result that used it, or you specifically need the ''parent_id'' tree or the per-language policy URLs and will re-verify each owner you rely on |
 | **WHOIS / RDAP** | current, but see the trap below | as a **fallback** for domains no list covers, and only with privacy-proxy filtering | | **WHOIS / RDAP** | current, but see the trap below | as a **fallback** for domains no list covers, and only with privacy-proxy filtering |
 | **TLS certificates, DNS/SOA, CNAME chains** | current | first-party CDN and sibling-domain detection, which the ownership lists are worst at: {[steffens2021_blockparty]} needed it because "those lists frequently miss connections among two hostnames, e.g., ''twitch.tv'' and ''twitchcdn.net''" | | **TLS certificates, DNS/SOA, CNAME chains** | current | first-party CDN and sibling-domain detection, which the ownership lists are worst at: {[steffens2021_blockparty]} needed it because "those lists frequently miss connections among two hostnames, e.g., ''twitch.tv'' and ''twitchcdn.net''" |
Line 300: Line 325:
  
 {[utz2023_rarely]} is worth reading before you pick one: it compared five third-party categorisations, including WhoTracks.me and Tracker Radar, and reports that "categorizations differ in granularity and focus" while overlapping substantially — the granularity axis again, stated from inside the literature. {[utz2023_rarely]} is worth reading before you pick one: it compared five third-party categorisations, including WhoTracks.me and Tracker Radar, and reports that "categorizations differ in granularity and focus" while overlapping substantially — the granularity axis again, stated from inside the literature.
 +
 +===== Assembling the pipeline =====
 +
 +The step most often missing from a paper is not the lookup, it is what surrounds it. To turn a request log into "//N//% of sites contact Google":
 +
 +  - **Extract the request host**, and keep the visited site's host beside it.
 +  - **Fold both to a registrable domain** using the **ICANN section** of a dated PSL. Not the private section — see the trap above.
 +  - **Resolve the request domain to an owner**, walking parent labels rather than looking up the exact key, and record which key matched.
 +  - **Resolve the visited site to an owner too, and drop the request if they are the same owner.** This is the step that is almost always silently skipped, and it is what the ownership list is //for//: {[wu2025_appprivacyreport]} describes webXray's list as "allowing distinction between first- and third-party domains". Without it, ''google.com'' embedding ''gstatic.com'' counts as third-party tracking, and any site whose CDN is on a sibling domain is over-counted. eTLD+1 comparison alone does not do it — that is exactly why {[steffens2021_blockparty]} needed eight person-hours of hand-vetting.
 +  - **Aggregate per site, not per request**, and state the unit: "sites with at least one request to an owner" is not "requests", is not "domains", and is not "owners".
 +  - **Report the unattributed remainder** twice: as a share of domains and as a share of sites or requests. Those differ by more than an order of magnitude here.
 +
 +Every one of those six steps is a place where two papers measuring "the same thing" diverge, and only the third is about which list you picked.
  
 ===== What to report in a paper ===== ===== What to report in a paper =====
Line 309: Line 347:
   * **the granularity you resolved to** — brand, operating legal entity, or ultimate parent — and, if you walked a hierarchy, how deep and how you handled cycles and missing parents.   * **the granularity you resolved to** — brand, operating legal entity, or ultimate parent — and, if you walked a hierarchy, how deep and how you handled cycles and missing parents.
   * **the fallback order**, if you merged sources, and the conflict rule. {[sanchezrola2021_journey]} and {[yang2020_comparative]} both used strict priority orders and both said so; that is the standard to meet.   * **the fallback order**, if you merged sources, and the conflict rule. {[sanchezrola2021_journey]} and {[yang2020_comparative]} both used strict priority orders and both said so; that is the standard to meet.
-  * **how many domains you could not attribute**, as a count and as a share of //requests or sites//, not just of domains. The two differ enormously: a list can cover 1.4% of domains and 54.5% of prevalence weight. Say which denominator your unattributed share uses.+  * **how many domains you could not attribute**, as a count and as a share of //requests or sites//, not just of domains. The two differ enormously: a list can cover 2.1% of domains and 58.6% of prevalence weight. Say which denominator your unattributed share uses.
   * **your privacy-proxy filter**, if WHOIS was involved, and the list of proxy strings you removed.   * **your privacy-proxy filter**, if WHOIS was involved, and the list of proxy strings you removed.
   * **the eTLD+1 rule and the PSL version**, since every one of these lists keys on the registrable domain and the PSL changes. Do not ship a frozen copy {[mcquistin2023_psl]}.   * **the eTLD+1 rule and the PSL version**, since every one of these lists keys on the registrable domain and the PSL changes. Do not ship a frozen copy {[mcquistin2023_psl]}.
programming/crawler/webxray.1786953255.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki