User Tools

Site Tools


statistics

Statistics

This namespace is for what these methods do to web-measurement data — a sample that is a ranking list, a site that shares a tag manager with two hundred others, a crawl that generates hypotheses the way it generates rows, a hand-coded ground truth, a plan deposited before the data. It is not a statistics textbook. A general account of a t-test, a logit, or Cohen's kappa belongs in one; what belongs here is the unit-of-analysis error, the family that a crawl invents, and the reporting gaps this corpus actually has. The publication corpus behind these pages is seven venues (CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P, 2010–2026, 5,859 extracted papers). 1,762 ran statistical inference; 1,357 recruited human participants; 3,318 hand-coded something. Each child names its own population.

A namespace page outlines the pages inside it rather than carrying its own content. 1) All 6 children below are written. start used to list them inline without a namespace landing; this page is that landing.

The pages

Page What a student needs it for
Hypothesis testing Which tests this field uses, and the unit-of-analysis problem when the unit is a site.
Pvalue corrections Multiplicity as a property of the crawl, not a decision you made.
Regression Modelling an outcome while accounting for covariates and for dependence between rows.
Biases Selection, survivorship, vantage and denominator bias as they appear in a web measurement.
Interrater agreement Reliability of hand-coded ground truth, and what to report.
Study preregistration Depositing the analysis plan before the data; the homograph with pre-registered domains.

Hypothesis testing and p-value corrections share a population on purpose (papers that ran a test), so they compose. Regression is the “by how much, holding other things constant” question; the child is about the forms that takes on web-measurement rows. Biases is the page that names what the Design choices did to the number. Inter-rater agreement is for the slice that stopped being automatic. Preregistration is the plan, not the artifact (Artifacts) and not the ethics review (Ethics).

Where this namespace stops

  • Sampling / Website selection / Crawling location / Longitudinal — choosing the frame, the vantage, the pin. Biases measures what those choices cost; it does not choose them.
  • User studies — recruiting people. Regression and preregistration are often user-study methods that crawl papers borrow.
  • Literature review — the keyword-derived denominator. A related-work count is not a statistical estimator, but it has the same missing-denominator failure mode.
  • Public relations — the sentence that travels. The denominator has to be in it.

Methodology and limitations of these figures

The 5,859 / 1,762 / 1,357 / 3,318 are paper counts from the 5,859-paper extraction (seven venues, 2010–2026). 2025–2026 venue-years are provisional — see corpus. The 6 is a wiki-page count as of 2026-08-27. Queries: statistics. Joint sitting: design.

1)
contributing, “Namespace and page structure”.
You could leave a comment if you were logged in.
statistics.txt · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki