User Tools

Site Tools


statistics:biases

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Next revision
Previous revision
statistics:biases [2026/08/20 20:13] – New page: selection bias from top lists, survivorship in repeated crawls, vantage-point bias, denominator bias. Figures from the 5,859-paper extraction with per-question denominators; 32 quotes verified. Authored by Claude karel.kubicek.claudestatistics:biases [2026/08/21 14:50] (current) – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude
Line 179: Line 179:
 | Reporting a **single landing-page-only** figure as a site-level rate | **Superseded** where subsites are reachable. | {[urban2020beyond]}{[aqeel2020_landing]} | | Reporting a **single landing-page-only** figure as a site-level rate | **Superseded** where subsites are reachable. | {[urban2020beyond]}{[aqeel2020_landing]} |
  
-Two things this table cannot tell you, both because the 2025–2026 slice is the thinnest in the corpus (see [[#Worked Example]]). **No bias characterisation specific to LLM-based classification appears** in the corpus's 2025–2026 material, although LLM classification itself is now common there: **166 of the 1,185 papers from 2025–2026 name a language model in a classification tuple (14.0%)**, and only 4 of those 166 mention bias or calibration anywhere — none of them about the selection or calibration behaviour of LLM labelling on web measurement data. So if your pipeline puts a model in the labelling step, its bias is an open question here, not a settled one. And **no paper in this corpus characterises which sites drop out of a crawl and how that biases the result**; it is listed as an open question on [[Design:Sampling#Open Questions]] and remains one.+Two things this table cannot tell you, both because the 2025–2026 slice is the thinnest in the corpus (see [[#worked_examplethe_biases_of_this_site_s_own_corpus|the worked example]]). **No bias characterisation specific to LLM-based classification appears** in the corpus's 2025–2026 material, although LLM classification itself is now common there: **166 of the 1,185 papers from 2025–2026 name a language model in a classification tuple (14.0%)**, and only 4 of those 166 mention bias or calibration anywhere — none of them about the selection or calibration behaviour of LLM labelling on web measurement data. So if your pipeline puts a model in the labelling step, its bias is an open question here, not a settled one. And **no paper in this corpus characterises which sites drop out of a crawl and how that biases the result**; it is listed as an open question on [[Design:Sampling#Open Questions]] and remains one.
  
 ===== Use in Publications ===== ===== Use in Publications =====
Line 374: Line 374:
 ===== Open Questions ===== ===== Open Questions =====
  
-  * <wrap todo>**The drop-out set has not been characterised.** No paper in this corpus measures which sites fail to load in a crawl and how their composition differs from the sites that load — the measurement that would let every crawling paper bound its own survivorship. An external sweep found one adjacent short paper on crawl refusals in Common Crawl, which is about server-side blocking generally rather than about how the lost set differs, so the question stands. Also open on [[Design:Sampling#Open Questions]].</wrap> +<WRAP todo> 
-  * <wrap todo>**No bias analysis for LLM-based labelling inside the seven venues.** LLM classification appears in the 2025–2026 corpus slice (166 of 1,185 papers name a model in a classification tuple) and a characterisation of its selection and calibration behaviour does not. The closest work is {[bozzolan2026_llmweb]}, a 2026 arXiv preprint that benchmarks LLM website classification for exactly this purpose — it establishes that model and configuration choice matters a great deal, which is the premise of the question rather than its answer. What is missing is a per-class error and calibration profile a measurement paper could cite to bound its own labelling bias. If you know of one, please add it.</wrap> +  * **The drop-out set has not been characterised.** No paper in this corpus measures which sites fail to load in a crawl and how their composition differs from the sites that load — the measurement that would let every crawling paper bound its own survivorship. An external sweep found one adjacent short paper on crawl refusals in Common Crawl, which is about server-side blocking generally rather than about how the lost set differs, so the question stands. Also open on [[Design:Sampling#Open Questions]]. 
-  * <wrap todo>**No post-2022 replacement for the top-list-against-ground-truth comparison.** Ruth et al. {[ruth2022_toppling]} used Cloudflare resolver data in February 2022; CrUX has since been folded into Tranco's default list, so the thing they measured no longer exists in the same form and the comparison has not been re-run.</wrap> +  * **No bias analysis for LLM-based labelling inside the seven venues.** LLM classification appears in the 2025–2026 corpus slice (166 of 1,185 papers name a model in a classification tuple) and a characterisation of its selection and calibration behaviour does not. The closest work is {[bozzolan2026_llmweb]}, a 2026 arXiv preprint that benchmarks LLM website classification for exactly this purpose — it establishes that model and configuration choice matters a great deal, which is the premise of the question rather than its answer. What is missing is a per-class error and calibration profile a measurement paper could cite to bound its own labelling bias. If you know of one, please add it. 
-  * <wrap todo>**Attrition vocabulary.** The probe on this page finds phrases. A hand-coded sample of repeated-crawl papers would establish how many report attrition //numerically// without using any of the vocabulary — the figure that would turn this page's upper bounds into estimates.</wrap>+  * **No post-2022 replacement for the top-list-against-ground-truth comparison.** Ruth et al. {[ruth2022_toppling]} used Cloudflare resolver data in February 2022; CrUX has since been folded into Tranco's default list, so the thing they measured no longer exists in the same form and the comparison has not been re-run. 
 +  * **Attrition vocabulary.** The probe on this page finds phrases. A hand-coded sample of repeated-crawl papers would establish how many report attrition //numerically// without using any of the vocabulary — the figure that would turn this page's upper bounds into estimates. 
 +</WRAP>
  
 ===== Related Pages ===== ===== Related Pages =====
statistics/biases.1787256834.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki