User Tools

Site Tools


statistics:biases

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
statistics:biases [2026/08/20 20:15] – Fix a broken same-page anchor to the worked example. Authored by Claude karel.kubicek.claudestatistics:biases [2026/08/21 14:50] (current) – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude
Line 374: Line 374:
 ===== Open Questions ===== ===== Open Questions =====
  
-  * <wrap todo>**The drop-out set has not been characterised.** No paper in this corpus measures which sites fail to load in a crawl and how their composition differs from the sites that load — the measurement that would let every crawling paper bound its own survivorship. An external sweep found one adjacent short paper on crawl refusals in Common Crawl, which is about server-side blocking generally rather than about how the lost set differs, so the question stands. Also open on [[Design:Sampling#Open Questions]].</wrap> +<WRAP todo> 
-  * <wrap todo>**No bias analysis for LLM-based labelling inside the seven venues.** LLM classification appears in the 2025–2026 corpus slice (166 of 1,185 papers name a model in a classification tuple) and a characterisation of its selection and calibration behaviour does not. The closest work is {[bozzolan2026_llmweb]}, a 2026 arXiv preprint that benchmarks LLM website classification for exactly this purpose — it establishes that model and configuration choice matters a great deal, which is the premise of the question rather than its answer. What is missing is a per-class error and calibration profile a measurement paper could cite to bound its own labelling bias. If you know of one, please add it.</wrap> +  * **The drop-out set has not been characterised.** No paper in this corpus measures which sites fail to load in a crawl and how their composition differs from the sites that load — the measurement that would let every crawling paper bound its own survivorship. An external sweep found one adjacent short paper on crawl refusals in Common Crawl, which is about server-side blocking generally rather than about how the lost set differs, so the question stands. Also open on [[Design:Sampling#Open Questions]]. 
-  * <wrap todo>**No post-2022 replacement for the top-list-against-ground-truth comparison.** Ruth et al. {[ruth2022_toppling]} used Cloudflare resolver data in February 2022; CrUX has since been folded into Tranco's default list, so the thing they measured no longer exists in the same form and the comparison has not been re-run.</wrap> +  * **No bias analysis for LLM-based labelling inside the seven venues.** LLM classification appears in the 2025–2026 corpus slice (166 of 1,185 papers name a model in a classification tuple) and a characterisation of its selection and calibration behaviour does not. The closest work is {[bozzolan2026_llmweb]}, a 2026 arXiv preprint that benchmarks LLM website classification for exactly this purpose — it establishes that model and configuration choice matters a great deal, which is the premise of the question rather than its answer. What is missing is a per-class error and calibration profile a measurement paper could cite to bound its own labelling bias. If you know of one, please add it. 
-  * <wrap todo>**Attrition vocabulary.** The probe on this page finds phrases. A hand-coded sample of repeated-crawl papers would establish how many report attrition //numerically// without using any of the vocabulary — the figure that would turn this page's upper bounds into estimates.</wrap>+  * **No post-2022 replacement for the top-list-against-ground-truth comparison.** Ruth et al. {[ruth2022_toppling]} used Cloudflare resolver data in February 2022; CrUX has since been folded into Tranco's default list, so the thing they measured no longer exists in the same form and the comparison has not been re-run. 
 +  * **Attrition vocabulary.** The probe on this page finds phrases. A hand-coded sample of repeated-crawl papers would establish how many report attrition //numerically// without using any of the vocabulary — the figure that would turn this page's upper bounds into estimates. 
 +</WRAP>
  
 ===== Related Pages ===== ===== Related Pages =====
statistics/biases.1787256939.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki