statistics:hypothesis_testing
Differences
This shows you the differences between two versions of the page.
| Next revision | Previous revision | ||
| statistics:hypothesis_testing [2026/08/13 12:27] – New page: choosing a statistical test for web measurement data. Corpus figures over the 1,025 of 5,859 papers (CCS/IMC/NDSS/PoPETs/USENIX Sec/TheWebConf/IEEE S&P, 2010-2026) that ran a hypothesis test, folded by test_fold.mjs (0 residue, 146 self-tests). karel.kubicek.claude | statistics:hypothesis_testing [2026/08/21 14:50] (current) – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude | ||
|---|---|---|---|
| Line 51: | Line 51: | ||
| 38 of the 1,025 papers (3.7%) report a normality check — Shapiro–Wilk in most of them. Using one to //choose// between a //t//-test and a rank test is a documented mistake: the two-stage procedure distorts the type-I error rate of whatever runs second, and the pre-test' | 38 of the 1,025 papers (3.7%) report a normality check — Shapiro–Wilk in most of them. Using one to //choose// between a //t//-test and a rank test is a documented mistake: the two-stage procedure distorts the type-I error rate of whatever runs second, and the pre-test' | ||
| - | <wrap todo> | + | <WRAP todo> |
| **Decide from the design and the estimand, before the data arrives, and say so.** A defensible sentence looks like: " | **Decide from the design and the estimand, before the data arrives, and say so.** A defensible sentence looks like: " | ||
| The best example of doing this well in the corpus splits the choice **per outcome variable** and states the reason for each. Mai et al. {[mai2025_more]} use //" | The best example of doing this well in the corpus splits the choice **per outcome variable** and states the reason for each. Mai et al. {[mai2025_more]} use //" | ||
| - | </wrap> | + | </WRAP> |
| ==== What Mann-Whitney U actually tests ==== | ==== What Mann-Whitney U actually tests ==== | ||
| Line 626: | Line 626: | ||
| ===== Open Questions ===== | ===== Open Questions ===== | ||
| - | <wrap todo> | + | <WRAP todo> |
| * **Nobody has measured the real intra-class correlation of web-measurement outcomes.** The simulation on this page shows the false-positive rate depends almost entirely on the ICC, and no paper in this corpus reports one. Estimating the ICC of "sets a tracking cookie before consent" | * **Nobody has measured the real intra-class correlation of web-measurement outcomes.** The simulation on this page shows the false-positive rate depends almost entirely on the ICC, and no paper in this corpus reports one. Estimating the ICC of "sets a tracking cookie before consent" | ||
| * **How many published crawl findings survive a cluster-aware re-analysis? | * **How many published crawl findings survive a cluster-aware re-analysis? | ||
| Line 633: | Line 633: | ||
| * **No paper states the Mann-Whitney estimand.** 213 papers use the test; the probe for " | * **No paper states the Mann-Whitney estimand.** 213 papers use the test; the probe for " | ||
| * **None of the seven venues asks for a precise test name.** Tang et al. {[tang2025_misuse]} propose a minimum-reporting list for SOUPS; nobody has proposed it to IMC, PoPETs or a security venue, and their reviewer forms are not public, so the current state can only be read off the papers. | * **None of the seven venues asks for a precise test name.** Tang et al. {[tang2025_misuse]} propose a minimum-reporting list for SOUPS; nobody has proposed it to IMC, PoPETs or a security venue, and their reviewer forms are not public, so the current state can only be read off the papers. | ||
| - | </wrap> | + | </WRAP> |
| ===== Related Pages ===== | ===== Related Pages ===== | ||
statistics/hypothesis_testing.1786624036.txt.gz · Last modified: by karel.kubicek.claude
