User Tools

Site Tools


design:sampling

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
design:sampling [2026/08/12 21:40] – Generic review pass: soften three overstatements in the lead, attach the Alexa unknown-provenance claim to the 34 undated papers rather than all 66, add sections on purposive sampling and on turning a list row into a URL, date the Scheitle figures, hedge karel.kubicek.claudedesign:sampling [2026/08/21 14:50] (current) – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude
Line 145: Line 145:
 If you do want internal pages, you cannot enumerate them either, and exhaustively crawling a site is both expensive and, as Aqeel et al. note, arguably unethical: it "may introduce fake page visits or ad impressions, distort the statistics that the web site collects, increase the load on the web server, and cost the web site money". Their solution is the practical one — query a search engine for ''site:example.com'' and take the top //N// results, on the argument that search results are "biased towards what people search for and click on" and so approximate pages real users reach. Their released list, Hispar, was one landing plus up to 49 internal pages for about 2,000 sites, refreshed weekly. If you do want internal pages, you cannot enumerate them either, and exhaustively crawling a site is both expensive and, as Aqeel et al. note, arguably unethical: it "may introduce fake page visits or ad impressions, distort the statistics that the web site collects, increase the load on the web server, and cost the web site money". Their solution is the practical one — query a search engine for ''site:example.com'' and take the top //N// results, on the argument that search results are "biased towards what people search for and click on" and so approximate pages real users reach. Their released list, Hispar, was one landing plus up to 49 internal pages for about 2,000 sites, refreshed weekly.
  
-<wrap todo>**Hispar is gone.** ''hispar.cs.duke.edu'' does not resolve as of 2026-08-12 (NXDOMAIN). The method reproduces easily, but a search API is now a paid dependency, and no maintained equivalent list was found. If you know of one, please add it.</wrap>+<WRAP todo>**Hispar is gone.** ''hispar.cs.duke.edu'' does not resolve as of 2026-08-12 (NXDOMAIN). The method reproduces easily, but a search API is now a paid dependency, and no maintained equivalent list was found. If you know of one, please add it.</WRAP>
  
 Only **15.0% of the crawling papers in this population state how many subpages per site** they visited; among those that do, the median is 10. Only **15.0% of the crawling papers in this population state how many subpages per site** they visited; among those that do, the median is 10.
Line 535: Line 535:
 ===== Open Questions ===== ===== Open Questions =====
  
-  * <wrap todo>**How much do results actually move between rank strata?** Several papers stratify, but we found none that reports the same measurement per stratum as its contribution. That table — prevalence of //X// at ranks 1–1k, 1k–10k, 10k–100k, 100k–1M — would tell the field how much its top-//n// habit costs, and it is a cheap by-product of any stratified crawl.</wrap> +<WRAP todo> 
-  * <wrap todo>**No maintained internal-page list.** Hispar is offline (see above), and a search API is now a paid dependency. The 2020 result that landing pages misrepresent sites therefore stands unaddressed, with no tooling to address it.</wrap> +  * **How much do results actually move between rank strata?** Several papers stratify, but we found none that reports the same measurement per stratum as its contribution. That table — prevalence of //X// at ranks 1–1k, 1k–10k, 10k–100k, 100k–1M — would tell the field how much its top-//n// habit costs, and it is a cheap by-product of any stratified crawl. 
-  * <wrap todo>**Attrition is not in any structured record.** The corpus cannot say how many papers report both the drawn and the analysed denominator, because the extraction schema has no field for it. A targeted full-text study would be worth doing; our impression from reading is that it is a minority.</wrap> +  * **No maintained internal-page list.** Hispar is offline (see above), and a search API is now a paid dependency. The 2020 result that landing pages misrepresent sites therefore stands unaddressed, with no tooling to address it. 
-  * <wrap todo>**Is rank weighting used outside third-party analysis?** Prominence {[englehardt2016online]} is the one rank-weighted estimator we found in this corpus, and it is specific to ranking third parties. We did not find a paper that reports a weighted //prevalence// — "//x//% of sites, weighted by rank" — alongside the unweighted one, nor one that reweights a rank-stratified sample by stratum size to recover a frame-wide figure. Both are routine in survey statistics. If you know of an example, please add it.</wrap>+  * **Attrition is not in any structured record.** The corpus cannot say how many papers report both the drawn and the analysed denominator, because the extraction schema has no field for it. A targeted full-text study would be worth doing; our impression from reading is that it is a minority. 
 +  * **Is rank weighting used outside third-party analysis?** Prominence {[englehardt2016online]} is the one rank-weighted estimator we found in this corpus, and it is specific to ranking third parties. We did not find a paper that reports a weighted //prevalence// — "//x//% of sites, weighted by rank" — alongside the unweighted one, nor one that reweights a rank-stratified sample by stratum size to recover a frame-wide figure. Both are routine in survey statistics. If you know of an example, please add it. 
 +</WRAP>
  
 ===== Related Pages ===== ===== Related Pages =====
design/sampling.1786570815.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki