User Tools

Site Tools


statistics:hypothesis_testing

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
statistics:hypothesis_testing [2026/08/13 12:54] – [Open Questions] formatting adminstatistics:hypothesis_testing [2026/08/21 14:50] (current) – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude
Line 51: Line 51:
 38 of the 1,025 papers (3.7%) report a normality check — Shapiro–Wilk in most of them. Using one to //choose// between a //t//-test and a rank test is a documented mistake: the two-stage procedure distorts the type-I error rate of whatever runs second, and the pre-test's own power depends on //n// in the wrong direction, so it waves through non-normality in small samples and rejects trivial non-normality in large ones. Rochon, Gondan & Kieser {[rochon2012_totest]} and Rasch, Kubinger & Moder {[rasch2011_pretesting]} both conclude that pre-testing does not pay off. 38 of the 1,025 papers (3.7%) report a normality check — Shapiro–Wilk in most of them. Using one to //choose// between a //t//-test and a rank test is a documented mistake: the two-stage procedure distorts the type-I error rate of whatever runs second, and the pre-test's own power depends on //n// in the wrong direction, so it waves through non-normality in small samples and rejects trivial non-normality in large ones. Rochon, Gondan & Kieser {[rochon2012_totest]} and Rasch, Kubinger & Moder {[rasch2011_pretesting]} both conclude that pre-testing does not pay off.
  
-<wrap todo>+<WRAP todo>
 **Decide from the design and the estimand, before the data arrives, and say so.** A defensible sentence looks like: "Because our outcome is a count with a long right tail and we care about the typical site rather than the mean, we pre-specified rank-based tests." An indefensible one is "Shapiro–Wilk was significant, so we used Mann-Whitney." **Decide from the design and the estimand, before the data arrives, and say so.** A defensible sentence looks like: "Because our outcome is a count with a long right tail and we care about the typical site rather than the mean, we pre-specified rank-based tests." An indefensible one is "Shapiro–Wilk was significant, so we used Mann-Whitney."
  
 The best example of doing this well in the corpus splits the choice **per outcome variable** and states the reason for each. Mai et al. {[mai2025_more]} use //"one-way repeated measures ANOVA as the omnibus test since the ad load data follow a normal distribution"// for one outcome and, //"for the hypothesis on predatory ad rates, since the rates do not approximately follow a normal distribution, we use Friedman test as the omnibus test"// for another, in the same paper.((Both sentences are in the paper's PDF and in **neither** of the corpus's stored text renderings of it — see [[provenance:statistics:hypothesis_testing]] for that discrepancy, which is a fact about the corpus and not about the paper.)) The best example of doing this well in the corpus splits the choice **per outcome variable** and states the reason for each. Mai et al. {[mai2025_more]} use //"one-way repeated measures ANOVA as the omnibus test since the ad load data follow a normal distribution"// for one outcome and, //"for the hypothesis on predatory ad rates, since the rates do not approximately follow a normal distribution, we use Friedman test as the omnibus test"// for another, in the same paper.((Both sentences are in the paper's PDF and in **neither** of the corpus's stored text renderings of it — see [[provenance:statistics:hypothesis_testing]] for that discrepancy, which is a fact about the corpus and not about the paper.))
-</wrap>+</WRAP>
  
 ==== What Mann-Whitney U actually tests ==== ==== What Mann-Whitney U actually tests ====
Line 626: Line 626:
 ===== Open Questions ===== ===== Open Questions =====
  
 +<WRAP todo>
   * **Nobody has measured the real intra-class correlation of web-measurement outcomes.** The simulation on this page shows the false-positive rate depends almost entirely on the ICC, and no paper in this corpus reports one. Estimating the ICC of "sets a tracking cookie before consent" by tag manager, by CMS and by hosting provider is a small, self-contained, immediately useful study, and it would tell every crawl paper how badly its //p//-values are wrong.   * **Nobody has measured the real intra-class correlation of web-measurement outcomes.** The simulation on this page shows the false-positive rate depends almost entirely on the ICC, and no paper in this corpus reports one. Estimating the ICC of "sets a tracking cookie before consent" by tag manager, by CMS and by hosting provider is a small, self-contained, immediately useful study, and it would tell every crawl paper how badly its //p//-values are wrong.
   * **How many published crawl findings survive a cluster-aware re-analysis?** Answerable on any paper that released per-site data (see [[:Artifacts]]), and nobody has done it. The three corpus papers that cluster are all 2023+, so essentially the whole literature is un-re-analysed.   * **How many published crawl findings survive a cluster-aware re-analysis?** Answerable on any paper that released per-site data (see [[:Artifacts]]), and nobody has done it. The three corpus papers that cluster are all 2023+, so essentially the whole literature is un-re-analysed.
Line 632: Line 633:
   * **No paper states the Mann-Whitney estimand.** 213 papers use the test; the probe for "stochastic superiority" or "stochastic dominance" in that sense returns zero. Whether authors know and do not write it, or write "median" because they believe it, is not answerable from text.   * **No paper states the Mann-Whitney estimand.** 213 papers use the test; the probe for "stochastic superiority" or "stochastic dominance" in that sense returns zero. Whether authors know and do not write it, or write "median" because they believe it, is not answerable from text.
   * **None of the seven venues asks for a precise test name.** Tang et al. {[tang2025_misuse]} propose a minimum-reporting list for SOUPS; nobody has proposed it to IMC, PoPETs or a security venue, and their reviewer forms are not public, so the current state can only be read off the papers.   * **None of the seven venues asks for a precise test name.** Tang et al. {[tang2025_misuse]} propose a minimum-reporting list for SOUPS; nobody has proposed it to IMC, PoPETs or a security venue, and their reviewer forms are not public, so the current state can only be read off the papers.
 +</WRAP>
 ===== Related Pages ===== ===== Related Pages =====
  
statistics/hypothesis_testing.1786625662.txt.gz · Last modified: by admin

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki