| Both sides previous revisionPrevious revision | |
| statistics:hypothesis_testing [2026/08/21 08:35] – [Open Questions] karelkubicek | statistics:hypothesis_testing [2026/08/21 14:50] (current) – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude |
|---|
| 38 of the 1,025 papers (3.7%) report a normality check — Shapiro–Wilk in most of them. Using one to //choose// between a //t//-test and a rank test is a documented mistake: the two-stage procedure distorts the type-I error rate of whatever runs second, and the pre-test's own power depends on //n// in the wrong direction, so it waves through non-normality in small samples and rejects trivial non-normality in large ones. Rochon, Gondan & Kieser {[rochon2012_totest]} and Rasch, Kubinger & Moder {[rasch2011_pretesting]} both conclude that pre-testing does not pay off. | 38 of the 1,025 papers (3.7%) report a normality check — Shapiro–Wilk in most of them. Using one to //choose// between a //t//-test and a rank test is a documented mistake: the two-stage procedure distorts the type-I error rate of whatever runs second, and the pre-test's own power depends on //n// in the wrong direction, so it waves through non-normality in small samples and rejects trivial non-normality in large ones. Rochon, Gondan & Kieser {[rochon2012_totest]} and Rasch, Kubinger & Moder {[rasch2011_pretesting]} both conclude that pre-testing does not pay off. |
| |
| <wrap todo> | <WRAP todo> |
| **Decide from the design and the estimand, before the data arrives, and say so.** A defensible sentence looks like: "Because our outcome is a count with a long right tail and we care about the typical site rather than the mean, we pre-specified rank-based tests." An indefensible one is "Shapiro–Wilk was significant, so we used Mann-Whitney." | **Decide from the design and the estimand, before the data arrives, and say so.** A defensible sentence looks like: "Because our outcome is a count with a long right tail and we care about the typical site rather than the mean, we pre-specified rank-based tests." An indefensible one is "Shapiro–Wilk was significant, so we used Mann-Whitney." |
| |
| The best example of doing this well in the corpus splits the choice **per outcome variable** and states the reason for each. Mai et al. {[mai2025_more]} use //"one-way repeated measures ANOVA as the omnibus test since the ad load data follow a normal distribution"// for one outcome and, //"for the hypothesis on predatory ad rates, since the rates do not approximately follow a normal distribution, we use Friedman test as the omnibus test"// for another, in the same paper.((Both sentences are in the paper's PDF and in **neither** of the corpus's stored text renderings of it — see [[provenance:statistics:hypothesis_testing]] for that discrepancy, which is a fact about the corpus and not about the paper.)) | The best example of doing this well in the corpus splits the choice **per outcome variable** and states the reason for each. Mai et al. {[mai2025_more]} use //"one-way repeated measures ANOVA as the omnibus test since the ad load data follow a normal distribution"// for one outcome and, //"for the hypothesis on predatory ad rates, since the rates do not approximately follow a normal distribution, we use Friedman test as the omnibus test"// for another, in the same paper.((Both sentences are in the paper's PDF and in **neither** of the corpus's stored text renderings of it — see [[provenance:statistics:hypothesis_testing]] for that discrepancy, which is a fact about the corpus and not about the paper.)) |
| </wrap> | </WRAP> |
| |
| ==== What Mann-Whitney U actually tests ==== | ==== What Mann-Whitney U actually tests ==== |