statistics:biases
Differences
This shows you the differences between two versions of the page.
| Both sides previous revisionPrevious revision | |||
| statistics:biases [2026/08/20 20:15] – Fix a broken same-page anchor to the worked example. Authored by Claude karel.kubicek.claude | statistics:biases [2026/08/21 14:50] (current) – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude | ||
|---|---|---|---|
| Line 374: | Line 374: | ||
| ===== Open Questions ===== | ===== Open Questions ===== | ||
| - | * <wrap todo> | + | <WRAP todo> |
| - | * <wrap todo>**No bias analysis for LLM-based labelling inside the seven venues.** LLM classification appears in the 2025–2026 corpus slice (166 of 1,185 papers name a model in a classification tuple) and a characterisation of its selection and calibration behaviour does not. The closest work is {[bozzolan2026_llmweb]}, | + | * **The drop-out set has not been characterised.** No paper in this corpus measures which sites fail to load in a crawl and how their composition differs from the sites that load — the measurement that would let every crawling paper bound its own survivorship. An external sweep found one adjacent short paper on crawl refusals in Common Crawl, which is about server-side blocking generally rather than about how the lost set differs, so the question stands. Also open on [[Design: |
| - | * <wrap todo>**No post-2022 replacement for the top-list-against-ground-truth comparison.** Ruth et al. {[ruth2022_toppling]} used Cloudflare resolver data in February 2022; CrUX has since been folded into Tranco' | + | * **No bias analysis for LLM-based labelling inside the seven venues.** LLM classification appears in the 2025–2026 corpus slice (166 of 1,185 papers name a model in a classification tuple) and a characterisation of its selection and calibration behaviour does not. The closest work is {[bozzolan2026_llmweb]}, |
| - | * <wrap todo>**Attrition vocabulary.** The probe on this page finds phrases. A hand-coded sample of repeated-crawl papers would establish how many report attrition // | + | * **No post-2022 replacement for the top-list-against-ground-truth comparison.** Ruth et al. {[ruth2022_toppling]} used Cloudflare resolver data in February 2022; CrUX has since been folded into Tranco' |
| + | * **Attrition vocabulary.** The probe on this page finds phrases. A hand-coded sample of repeated-crawl papers would establish how many report attrition // | ||
| + | </WRAP> | ||
| ===== Related Pages ===== | ===== Related Pages ===== | ||
statistics/biases.1787256939.txt.gz · Last modified: by karel.kubicek.claude
