statistics:pvalue_corrections
Differences
This shows you the differences between two versions of the page.
| Both sides previous revisionPrevious revisionNext revision | Previous revision | ||
| statistics:pvalue_corrections [2026/08/13 06:56] – Review fixes: replace two false 'appears in zero' claims about forking paths (the string matches twice, both false positives); fix a footnote that said 42 exceeds 45; correct the per-year range to 0.0%-39.1% on 8-128; restore a plural in the Sunlight quot karel.kubicek.claude | statistics:pvalue_corrections [2026/08/21 14:50] (current) – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude | ||
|---|---|---|---|
| Line 6: | Line 6: | ||
| <WRAP important> | <WRAP important> | ||
| - | **Roughly a quarter of the papers that run a hypothesis test report a correction, and the papers with the most hypotheses | + | **Roughly a quarter of the papers that run a hypothesis test report a correction, and whether they do tracks whether they recruited people — not how many hypotheses |
| Of the 5,859 papers extracted from seven security, privacy and measurement venues (2010–2026), | Of the 5,859 papers extracted from seven security, privacy and measurement venues (2010–2026), | ||
| * **38.6%** of the 433 that recruited participants and ran no crawl | * **38.6%** of the 433 that recruited participants and ran no crawl | ||
| + | * **34.7%** of the 49 that did both | ||
| * **12.9%** of the 140 that ran a crawl and recruited nobody | * **12.9%** of the 140 that ran a crawl and recruited nobody | ||
| * **11.2%** of the 403 that did neither | * **11.2%** of the 403 that did neither | ||
| - | By venue, of each venue's hypothesis-test papers: **PoPETs 45.6%**, USENIX Security 28.5%, IEEE S&P 26.6%, CCS 20.9%, TheWebConf 14.4%, **IMC 14.1%**, **NDSS 10.8%**. Correction is a usable-privacy practice that measurement work has not picked up — and the crawl papers are the ones running 100+ tests: of the 189 papers that ran both a crawl and a hypothesis test, the median largest stated population is **38,090 units**, and **45.0% state a population of 100,000 or more**. | + | The two participant rows are the two high ones. **Correcting is a habit that arrived with human-subjects methodology, and a crawl paper that runs a hypothesis test corrects at a third the rate of a user study that does** — in the setting where the number of candidate hypotheses is set by the data rather than by a protocol. By venue: **PoPETs 45.6%**, USENIX Security 28.5%, IEEE S&P 26.6%, CCS 20.9%, TheWebConf 14.4%, **IMC 14.1%**, **NDSS 10.8%**. |
| - | **Bonferroni is still the modal choice in this field, including in 2025–2026, | + | **Bonferroni is still the modal choice in this field, including in 2025–2026, |
| - | And **almost nobody | + | And **hardly anyone |
| </ | </ | ||
| Line 41: | Line 42: | ||
| | **Per-treatment survival comparison over time** | arms × time points | Maass et al. {[maass2021_effective]}: | | **Per-treatment survival comparison over time** | arms × time points | Maass et al. {[maass2021_effective]}: | ||
| | **Per-violation-type compliance test** | one test per legal requirement per subgroup, and again per subgroup comparison | Bouhoula et al. {[bouhoula2024_automated]} apply Holm–Bonferroni twice: to pairwise violation comparisons across popularity ranks, and to Fisher' | | **Per-violation-type compliance test** | one test per legal requirement per subgroup, and again per subgroup comparison | Bouhoula et al. {[bouhoula2024_automated]} apply Holm–Bonferroni twice: to pairwise violation comparisons across popularity ranks, and to Fisher' | ||
| - | | **Per-regression-coefficient interaction term** | one test per coefficient, | + | | **Per-regression-coefficient interaction term** | one test per coefficient, |
| - | | **Per-site test** | one test per site. This is where //k// stops being tractable and the answer is a model, not a correction | no paper in this corpus | + | | **Per-site test** | one test per site. This is where //k// stops being tractable and the answer is a model, not a correction | none of the 269 papers with a correction tuple does this. Whether papers run per-site tests //without// correcting is not something |
| Two consequences a reviewer will pick up on. | Two consequences a reviewer will pick up on. | ||
| Line 49: | Line 50: | ||
| **The last row of the table is the honest one.** Once the natural unit of the test is the site, //k// is the size of the crawl, the correction is annihilating, | **The last row of the table is the honest one.** Once the natural unit of the test is the site, //k// is the size of the crawl, the correction is annihilating, | ||
| + | |||
| + | ===== Deciding What Counts as One Family ===== | ||
| + | |||
| + | This is the decision the procedures cannot make for you, it is the one a textbook does not cover, and it is where a crawl-shaped study goes wrong. Three consent arms × 40 countries is 120 cells. Is that one family of 120 tests, three families of 40, 40 families of 3, or — if you report every pairwise country contrast inside each arm — one family of 3 × 780 = 2,340? The answer changes every number in your results table, and **any of those answers is defensible if you fixed it before you looked.** None of them is defensible if you chose after. | ||
| + | |||
| + | The rule the corpus' | ||
| + | |||
| + | - **One family per confirmatory claim you will make in the abstract.** If the paper' | ||
| + | - **Then, per outcome variable.** Despres et al. {[despres2024_best]} fix α from //"a maximum of 25 hypotheses tested per outcome"//; | ||
| + | - **A model is one family, not one test per coefficient across the paper.** If you fit six regressions, | ||
| + | - **Everything else is exploratory.** Say so, report unadjusted // | ||
| + | - **Write the grouping down before the data arrive**, in the paper' | ||
| + | |||
| + | <WRAP important> | ||
| + | **One correction across the whole paper is the wrong answer, and it is the one people reach for.** It is not the conservative choice. It inflates //k// with tests nobody was going to interpret, which is exactly the cost the worked example below measures: the same twenty real effects go from eight reportable at //k// = 200 to one at //k// = 2,000, purely because the family got wider. **Splitting into stated families is not // | ||
| + | </ | ||
| ===== Which Procedure, and Which Are Current ===== | ===== Which Procedure, and Which Are Current ===== | ||
| Line 55: | Line 72: | ||
| ^ Procedure ^ Dated ^ Controls ^ Assumes about dependence ^ Papers ^ Verdict ^ | ^ Procedure ^ Dated ^ Controls ^ Assumes about dependence ^ Papers ^ Verdict ^ | ||
| - | | **Bonferroni** | 1936 / Dunn 1961 | FWER | nothing — valid under arbitrary dependence | **141 (52.4%)** | **Historical. Superseded | + | | **Bonferroni** | 1936 / Dunn 1961 | FWER | nothing — valid under arbitrary dependence | **141 (52.4%)** | **Historical |
| | **Holm** (step-down) | 1979 {[holm1979_sequentially]} | FWER | nothing | 66 (24.5%) | **Current. The default for a small confirmatory family.** Strictly dominates Bonferroni | | | **Holm** (step-down) | 1979 {[holm1979_sequentially]} | FWER | nothing | 66 (24.5%) | **Current. The default for a small confirmatory family.** Strictly dominates Bonferroni | | ||
| | **Šidák** | 1967 | FWER | independence | 3 (1.1%) | Marginal. Buys a fraction of a percent over Bonferroni and needs an assumption a crawl cannot support | | | **Šidák** | 1967 | FWER | independence | 3 (1.1%) | Marginal. Buys a fraction of a percent over Bonferroni and needs an assumption a crawl cannot support | | ||
| Line 70: | Line 87: | ||
| **What this field does:** Bonferroni remains the modal procedure in the most recent years available. Of the 45 hypothesis-test papers that corrected in 2025–2026, | **What this field does:** Bonferroni remains the modal procedure in the most recent years available. Of the 45 hypothesis-test papers that corrected in 2025–2026, | ||
| - | **What you should do:** use **Holm** for a small confirmatory family | + | **What you should do**, in this order: |
| + | |||
| + | - **Decide what the family is, and write it down, before you look at the data.** This matters more than which procedure you pick, and it is the one step no procedure can repair. See [[#Deciding What Counts as One Family]]. | ||
| + | - **Split confirmatory from exploratory**, | ||
| + | - **Then pick a procedure:** **Holm** for a small confirmatory family, **Benjamini–Hochberg** for a large exploratory one, **Benjamini–Yekutieli** instead of BH when you cannot argue the tests are positively dependent | ||
| + | - **Check the choice against your own numbers.** BY pays a factor of about ln(// | ||
| + | - **Report an effect size with an interval for every corrected comparison.** Past about 10,000 units per group the //p//-value stops being the binding constraint, and no correction fixes that. See [[#When Correcting Is the Wrong Lever]]. | ||
| **Nothing here has been overturned recently.** Multiplicity correction is one of the few methodological areas in web measurement where no 2020s development — LLM classification included — changed the answer. The procedures date from 1979, 1995 and 2001, and they are still the right ones. Where the field has moved is in //which// of them it reaches for, and slowly. Research-frontier work does exist — e-values and e-BH since 2022, selective inference — and **none of it has any presence in this corpus or any adopted standing**, so this page does not recommend it. | **Nothing here has been overturned recently.** Multiplicity correction is one of the few methodological areas in web measurement where no 2020s development — LLM classification included — changed the answer. The procedures date from 1979, 1995 and 2001, and they are still the right ones. Where the field has moved is in //which// of them it reaches for, and slowly. Research-frontier work does exist — e-values and e-BH since 2022, selective inference — and **none of it has any presence in this corpus or any adopted standing**, so this page does not recommend it. | ||
| - | What //has// moved is **reporting**, | + | What //has// moved is **reporting**, |
| </ | </ | ||
| Line 83: | Line 106: | ||
| * **A compliance or enforcement claim** — "this CMP violates Article 7(3)" — is FWER. One false accusation is a problem in itself, and the audience includes people who will read exactly one row. Correct so that the probability of //any// false claim is bounded: Holm. | * **A compliance or enforcement claim** — "this CMP violates Article 7(3)" — is FWER. One false accusation is a problem in itself, and the audience includes people who will read exactly one row. Correct so that the probability of //any// false claim is bounded: Holm. | ||
| * **A screen** — "which of these 400 third parties set an identifier before consent, so we can go look at them" — is FDR. You expect some of the shortlist to be wrong, you will follow up, and being unable to find anything is the real failure. Correct so that the //expected fraction// of wrong rows is bounded: Benjamini–Hochberg, | * **A screen** — "which of these 400 third parties set an identifier before consent, so we can go look at them" — is FDR. You expect some of the shortlist to be wrong, you will follow up, and being unable to find anything is the real failure. Correct so that the //expected fraction// of wrong rows is bounded: Benjamini–Hochberg, | ||
| - | * **A paper usually has both**, and the two families should be adjusted separately with the split stated. Nenadic et al. {[nenadic2026_overcoming]} do exactly | + | * **A paper usually has both**, and the two families should be adjusted separately with the split stated. Nenadic et al. {[nenadic2026_swiss]} do this for two contrast definitions of one question: BH //" |
| SciPy' | SciPy' | ||
| Line 332: | Line 355: | ||
| <WRAP important> | <WRAP important> | ||
| - | **At crawl scale, the binding constraint is almost never the // | + | **At crawl scale, the binding constraint is almost never the // |
| Correcting such a //p//-value is arithmetic performed on a quantity that stopped carrying information. What a reader needs is the **effect size and its interval**: the difference in percentage points, the risk ratio, the //r// or Cramér' | Correcting such a //p//-value is arithmetic performed on a quantity that stopped carrying information. What a reader needs is the **effect size and its interval**: the difference in percentage points, the risk ratio, the //r// or Cramér' | ||
| Line 351: | Line 374: | ||
| Each of those was a choice made with the data in front of you, and each is a test you effectively ran and did not report. Gelman & Loken {[gelman2014_statistical]} call this the garden of forking paths, and their point is precisely that it does not require any dishonesty — a single analysis path, chosen after seeing the data, has an uncontrolled error rate even though no multiple comparison was ever computed. | Each of those was a choice made with the data in front of you, and each is a test you effectively ran and did not report. Gelman & Loken {[gelman2014_statistical]} call this the garden of forking paths, and their point is precisely that it does not require any dishonesty — a single analysis path, chosen after seeing the data, has an uncontrolled error rate even though no multiple comparison was ever computed. | ||
| - | **No paper in this corpus engages with the idea.** The literal string //forking paths// matches two of the 5,869 full texts, and reading both shows they are packet-forwarding paths and symbolic-execution paths — see [[#What is missing entirely]]. The ASA statement on // | + | **No paper in this corpus engages with the idea.** The literal string //forking paths// matches two of the 5,869 full texts, and reading both shows they are packet-forwarding paths and symbolic-execution paths — see [[#What is missing entirely]]. The ASA statement on // |
| - | <wrap todo> | + | <WRAP todo> |
| - | The cheap version, if you will not preregister: | + | The cheap version, if you will not preregister: |
| - | </wrap> | + | </WRAP> |
| ==== Declining to correct is a legitimate choice, if you say so ==== | ==== Declining to correct is a legitimate choice, if you say so ==== | ||
| Line 424: | Line 447: | ||
| The last two rows are the reason the fold exists rather than a '' | The last two rows are the reason the fold exists rather than a '' | ||
| - | ==== It is rising, and Bonferroni | + | ==== It rose through the 2010s, then fell back — and Bonferroni |
| ^ Period ^ Ran a hypothesis test ^ Corrected ^ Share ^ Bonferroni ^ Holm ^ B–H ^ B–Y ^ | ^ Period ^ Ran a hypothesis test ^ Corrected ^ Share ^ Bonferroni ^ Holm ^ B–H ^ B–Y ^ | ||
| Line 433: | Line 456: | ||
| <WRAP important> | <WRAP important> | ||
| - | **Read the asterisk before | + | **The rise is solid; |
| + | |||
| + | The obvious excuse does not work. 2025–2026 are the provisional years — CCS 2026 and IMC 2026 have not been held, and IEEE S&P 2026 and TheWebConf 2026 abstracts are absent from OpenAlex | ||
| What is stable across the buckets is that **Bonferroni has never stopped being modal.** Every column above is restricted to hypothesis-test papers, so the four procedure counts and the " | What is stable across the buckets is that **Bonferroni has never stopped being modal.** Every column above is restricted to hypothesis-test papers, so the four procedure counts and the " | ||
| Line 442: | Line 467: | ||
| ==== Almost nobody states the size of the family ==== | ==== Almost nobody states the size of the family ==== | ||
| - | An adjusted //p// cannot be checked, reproduced or compared without // | + | An adjusted //p// cannot be checked, reproduced or compared without // |
| - | ^ What the paper states ^ Tuples ^ Share of 287 ^ | + | **Pass 1, over the 287 tuples that record a real correction** — what landed in '' |
| + | |||
| + | ^ What the tuple states ^ Tuples ^ Share of 287 ^ | ||
| | a number of comparisons, | | a number of comparisons, | ||
| | an adjusted α or significance threshold | 57 | 19.9% | | | an adjusted α or significance threshold | 57 | 19.9% | | ||
| | any reported detail at all ('' | | any reported detail at all ('' | ||
| + | |||
| + | **Pass 2, over the whole text of the 265 papers that corrected.** A tuple only carries one sentence, so this pass searches the full paper for any count of comparisons, | ||
| + | |||
| + | **So: 4.9% by tuple, 5.7% by paper. Roughly one corrected paper in twenty states the size of its family.** | ||
| The fourteen that do state it are the ones a reader can check, and they read like this: | The fourteen that do state it are the ones a reader can check, and they read like this: | ||
| Line 475: | Line 506: | ||
| | power-analysis | 72 | 7.0% | | | power-analysis | 72 | 7.0% | | ||
| | bayesian | 12 | 1.2% | | | bayesian | 12 | 1.2% | | ||
| - | | preregistration | 6 | 0.6% | | + | | preregistration |
| **69 papers report both an effect size and a correction. 178 report a correction and no effect size.** | **69 papers report both an effect size and a correction. 178 report a correction and no effect size.** | ||
| Line 501: | Line 532: | ||
| * **Two buckets in the fold are not corrections** and are excluded from every " | * **Two buckets in the fold are not corrections** and are excluded from every " | ||
| * **Sentinels are never answers.** '' | * **Sentinels are never answers.** '' | ||
| + | * **The family-size figure is measured twice and hand-checked.** The tuple pass sees one sentence per correction; the full-text pass sees the whole paper but needs guards against χ² notation, exponents and page furniture, and then a reading of all 20 surviving hits. 4.9% and 5.7% are the two answers, five rejected hits are named with reasons, and all 20 are printed with their sentences on the provenance page. **Neither pass is a bound**: a count in a distant table caption is missed by both. | ||
| * **A paper counts once**, never once per mention. 292 tuples across 269 papers. | * **A paper counts once**, never once per mention. 292 tuples across 269 papers. | ||
| * **'' | * **'' | ||
| Line 519: | Line 551: | ||
| - **The correction inside your power analysis**, if you are sizing a study rather than analysing one {[ho2025_efficacy]}. | - **The correction inside your power analysis**, if you are sizing a study rather than analysing one {[ho2025_efficacy]}. | ||
| - **If you decline to correct, say so and say why** — and then drop the word " | - **If you decline to correct, say so and say why** — and then drop the word " | ||
| - | - **Do not write " | + | - **For a test, write " |
| ===== Open Questions ===== | ===== Open Questions ===== | ||
| - | <wrap todo> | + | <WRAP todo> |
| * **No measurement paper in this corpus corrects a per-site family.** The design that most obviously creates thousands of hypotheses — one test per site — is never followed by a correction, so it is unknown whether the field considers that setting out of scope for testing, or simply does not test it. Both readings have consequences for how a reviewer should treat a per-site claim. | * **No measurement paper in this corpus corrects a per-site family.** The design that most obviously creates thousands of hypotheses — one test per site — is never followed by a correction, so it is unknown whether the field considers that setting out of scope for testing, or simply does not test it. Both readings have consequences for how a reviewer should treat a per-site claim. | ||
| * **65 papers fit a multilevel model and none frames it as a multiplicity strategy** — at least, none surfaced while reading correction passages, and no one has checked all 65. Whether partial pooling is already doing the work of a correction in this literature without anyone saying so is answerable and unanswered. | * **65 papers fit a multilevel model and none frames it as a multiplicity strategy** — at least, none surfaced while reading correction passages, and no one has checked all 65. Whether partial pooling is already doing the work of a correction in this literature without anyone saying so is answerable and unanswered. | ||
| * **Benjamini–Yekutieli is used by four papers**, and it is the procedure whose assumptions actually match site-level crawl data. Whether the BH results in this literature would survive BY is checkable on any paper that released its // | * **Benjamini–Yekutieli is used by four papers**, and it is the procedure whose assumptions actually match site-level crawl data. Whether the BH results in this literature would survive BY is checkable on any paper that released its // | ||
| * **The debate is absent.** Zero papers cite the ASA statement, Rothman, Perneger, or the garden of forking paths. A short SoK on inference practice in web measurement would be citing an empty shelf, which is unusual and useful. | * **The debate is absent.** Zero papers cite the ASA statement, Rothman, Perneger, or the garden of forking paths. A short SoK on inference practice in web measurement would be citing an empty shelf, which is unusual and useful. | ||
| - | * **No venue asks for //k//.** Reporting the family size is a one-line checklist item that would make every correction in the literature checkable, and no call for papers, | + | * **Nobody has asked these venues to require |
| * **The pipeline-level multiplicity is unmeasured.** Nobody has taken a published crawl, enumerated the defensible alternatives at each pipeline decision (seed list, rank cut, exclusion rule, classifier threshold), re-run all of them, and reported the spread of the headline figure. That is a multiverse analysis, it is straightforwardly fundable, and it would say more about this literature' | * **The pipeline-level multiplicity is unmeasured.** Nobody has taken a published crawl, enumerated the defensible alternatives at each pipeline decision (seed list, rank cut, exclusion rule, classifier threshold), re-run all of them, and reported the spread of the headline figure. That is a multiverse analysis, it is straightforwardly fundable, and it would say more about this literature' | ||
| - | </wrap> | + | </WRAP> |
| ===== Related Pages ===== | ===== Related Pages ===== | ||
| Line 540: | Line 572: | ||
| * [[Design: | * [[Design: | ||
| * [[Design: | * [[Design: | ||
| - | * [[Design: | + | * [[Design: |
| * [[: | * [[: | ||
| * [[provenance: | * [[provenance: | ||
statistics/pvalue_corrections.1786604174.txt.gz · Last modified: by karel.kubicek.claude
