User Tools

Site Tools


design:archives

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
design:archives [2026/08/13 22:44] – Extend: what an archive preserves and what it does not (static capture, escapes, coverage, vantage), corpus-backed Use in Publications, best-practice checklist, anachronism trap, self-recording alternative; correct five stale/wrong facts (Wayback size, Sa karel.kubicek.claudedesign:archives [2026/08/21 14:50] (current) – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude
Line 290: Line 290:
 ===== Open Questions ===== ===== Open Questions =====
  
-  * <wrap todo>**The archive comparison needs redoing, and the instrument for redoing it shrank.** Memento Time Travel was decommissioned in 2025; MemGator survives but reaches 13 archives rather than 30-plus, so Hantke et al.'s enumeration {[hantke2023you]} cannot be repeated as written and its 2022 coverage numbers are now four years old. Nobody has published a current census of which public web archives still crawl.</wrap> +<WRAP todo> 
-  * <wrap todo>**Nobody has re-measured static-vs-dynamic archive bias at scale.** The 73.9%/95.3% tracker gap {[hantke2023you]} rests on 2,026 sites, limited by rate-limiting. Zhu et al.'s {[zhu2025_toward]} 45.8% fidelity gap is measured on their own crawls, not on the Internet Archive's holdings. A large-scale measurement of what the Internet Archive itself is missing, per resource type, does not exist.</wrap> +  * **The archive comparison needs redoing, and the instrument for redoing it shrank.** Memento Time Travel was decommissioned in 2025; MemGator survives but reaches 13 archives rather than 30-plus, so Hantke et al.'s enumeration {[hantke2023you]} cannot be repeated as written and its 2022 coverage numbers are now four years old. Nobody has published a current census of which public web archives still crawl. 
-  * <wrap todo>**Monoculture risk is unquantified.** Roughly two thirds of archive-using papers depend on one archive, one crawler and one vantage point. What that shared bias does to the field's longitudinal results has not been studied — and Common Crawl, the only comparison Hantke et al. could make, was rejected on coverage rather than validated on agreement.</wrap> +  * **Nobody has re-measured static-vs-dynamic archive bias at scale.** The 73.9%/95.3% tracker gap {[hantke2023you]} rests on 2,026 sites, limited by rate-limiting. Zhu et al.'s {[zhu2025_toward]} 45.8% fidelity gap is measured on their own crawls, not on the Internet Archive's holdings. A large-scale measurement of what the Internet Archive itself is missing, per resource type, does not exist. 
-  * <wrap todo>**Archive-based measurement of regional phenomena has no instrument.** Both major sources crawl from the US. For GDPR-era consent, geo-blocking or regional advertising, there is no archive with an EU vantage point at usable coverage — the national archives fail the freshness test.</wrap>+  * **Monoculture risk is unquantified.** Roughly two thirds of archive-using papers depend on one archive, one crawler and one vantage point. What that shared bias does to the field's longitudinal results has not been studied — and Common Crawl, the only comparison Hantke et al. could make, was rejected on coverage rather than validated on agreement. 
 +  * **Archive-based measurement of regional phenomena has no instrument.** Both major sources crawl from the US. For GDPR-era consent, geo-blocking or regional advertising, there is no archive with an EU vantage point at usable coverage — the national archives fail the freshness test. 
 +</WRAP>
  
 ===== Related Pages ===== ===== Related Pages =====
design/archives.1786661095.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki