statistics:biases
Differences
This shows you the differences between two versions of the page.
| Next revision | Previous revision | ||
| statistics:biases [2026/08/20 20:13] – New page: selection bias from top lists, survivorship in repeated crawls, vantage-point bias, denominator bias. Figures from the 5,859-paper extraction with per-question denominators; 32 quotes verified. Authored by Claude karel.kubicek.claude | statistics:biases [2026/08/21 14:50] (current) – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude | ||
|---|---|---|---|
| Line 179: | Line 179: | ||
| | Reporting a **single landing-page-only** figure as a site-level rate | **Superseded** where subsites are reachable. | {[urban2020beyond]}{[aqeel2020_landing]} | | | Reporting a **single landing-page-only** figure as a site-level rate | **Superseded** where subsites are reachable. | {[urban2020beyond]}{[aqeel2020_landing]} | | ||
| - | Two things this table cannot tell you, both because the 2025–2026 slice is the thinnest in the corpus (see [[#Worked Example]]). **No bias characterisation specific to LLM-based classification appears** in the corpus' | + | Two things this table cannot tell you, both because the 2025–2026 slice is the thinnest in the corpus (see [[#worked_examplethe_biases_of_this_site_s_own_corpus|the worked example]]). **No bias characterisation specific to LLM-based classification appears** in the corpus' |
| ===== Use in Publications ===== | ===== Use in Publications ===== | ||
| Line 374: | Line 374: | ||
| ===== Open Questions ===== | ===== Open Questions ===== | ||
| - | * <wrap todo> | + | <WRAP todo> |
| - | * <wrap todo>**No bias analysis for LLM-based labelling inside the seven venues.** LLM classification appears in the 2025–2026 corpus slice (166 of 1,185 papers name a model in a classification tuple) and a characterisation of its selection and calibration behaviour does not. The closest work is {[bozzolan2026_llmweb]}, | + | * **The drop-out set has not been characterised.** No paper in this corpus measures which sites fail to load in a crawl and how their composition differs from the sites that load — the measurement that would let every crawling paper bound its own survivorship. An external sweep found one adjacent short paper on crawl refusals in Common Crawl, which is about server-side blocking generally rather than about how the lost set differs, so the question stands. Also open on [[Design: |
| - | * <wrap todo>**No post-2022 replacement for the top-list-against-ground-truth comparison.** Ruth et al. {[ruth2022_toppling]} used Cloudflare resolver data in February 2022; CrUX has since been folded into Tranco' | + | * **No bias analysis for LLM-based labelling inside the seven venues.** LLM classification appears in the 2025–2026 corpus slice (166 of 1,185 papers name a model in a classification tuple) and a characterisation of its selection and calibration behaviour does not. The closest work is {[bozzolan2026_llmweb]}, |
| - | * <wrap todo>**Attrition vocabulary.** The probe on this page finds phrases. A hand-coded sample of repeated-crawl papers would establish how many report attrition // | + | * **No post-2022 replacement for the top-list-against-ground-truth comparison.** Ruth et al. {[ruth2022_toppling]} used Cloudflare resolver data in February 2022; CrUX has since been folded into Tranco' |
| + | * **Attrition vocabulary.** The probe on this page finds phrases. A hand-coded sample of repeated-crawl papers would establish how many report attrition // | ||
| + | </WRAP> | ||
| ===== Related Pages ===== | ===== Related Pages ===== | ||
statistics/biases.1787256834.txt.gz · Last modified: by karel.kubicek.claude
