User Tools

Site Tools


statistics:pvalue_corrections

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
statistics:pvalue_corrections [2026/08/13 07:09] – Generic-review fixes (14 accepted): new 'Deciding What Counts as One Family' section; headline reframed to what the data shows (participants predict correcting, not family size) and the 'both 34.7%' cell restored; 'no remaining reason to use Bonferroni' r karel.kubicek.claudestatistics:pvalue_corrections [2026/08/21 14:50] (current) – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude
Line 19: Line 19:
 **Bonferroni is still the modal choice in this field, including in 2025–2026, and for testing there is no longer a reason to prefer it.** Of the 269 papers with a correction tuple, **141 (52.4%) use Bonferroni**, 66 (24.5%) Holm, 43 (16.0%) Benjamini–Hochberg, 4 (1.5%) Benjamini–Yekutieli. Holm {[holm1979_sequentially]} needs exactly the same assumptions as Bonferroni and can never reject fewer hypotheses — Datta et al. said so in this corpus in 2015 {[datta2015_automated]} — so //"we applied a Bonferroni correction"// should read //"we applied a Holm correction"// almost everywhere it is written. It survives for two things Holm does not do: **simultaneous confidence intervals**, and a **threshold you can pre-declare** before any //p//-value exists. See [[#Which Procedure, and Which Are Current]]. **Bonferroni is still the modal choice in this field, including in 2025–2026, and for testing there is no longer a reason to prefer it.** Of the 269 papers with a correction tuple, **141 (52.4%) use Bonferroni**, 66 (24.5%) Holm, 43 (16.0%) Benjamini–Hochberg, 4 (1.5%) Benjamini–Yekutieli. Holm {[holm1979_sequentially]} needs exactly the same assumptions as Bonferroni and can never reject fewer hypotheses — Datta et al. said so in this corpus in 2015 {[datta2015_automated]} — so //"we applied a Bonferroni correction"// should read //"we applied a Holm correction"// almost everywhere it is written. It survives for two things Holm does not do: **simultaneous confidence intervals**, and a **threshold you can pre-declare** before any //p//-value exists. See [[#Which Procedure, and Which Are Current]].
  
-And **hardly anyone says how many hypotheses were in the family.** 14 of 287 correction tuples (4.9%) put a count in the extracted method, detail or evidence sentence; widening the search to the whole paper, a count appears within ±1,500 characters of a procedure mention in **41 of 265 papers (15.5%)**. The true rate is between those two, and either way the large majority do not state //k// — without which the adjusted //p// is not reconstructible and the correction is not checkable.+And **hardly anyone says how many hypotheses were in the family: about one corrected paper in twenty.** Two independent measurements agree. 14 of 287 correction tuples (4.9%) put a count in the extracted method, detail or evidence sentence. Searching the whole paper instead — any count of comparisons, tests or hypotheses within ±1,500 characters of a procedure mention, then reading every hit — gives **15 of 265 papers (5.7%)**. Without //k// the adjusted //p// is not reconstructible and the correction is not checkable.
 </WRAP> </WRAP>
  
Line 42: Line 42:
 | **Per-treatment survival comparison over time** | arms × time points | Maass et al. {[maass2021_effective]}: //"a single Holm-Bonferroni correction [24] for all 45 significance tests"// | | **Per-treatment survival comparison over time** | arms × time points | Maass et al. {[maass2021_effective]}: //"a single Holm-Bonferroni correction [24] for all 45 significance tests"// |
 | **Per-violation-type compliance test** | one test per legal requirement per subgroup, and again per subgroup comparison | Bouhoula et al. {[bouhoula2024_automated]} apply Holm–Bonferroni twice: to pairwise violation comparisons across popularity ranks, and to Fisher's exact tests comparing each CMP's subset against all websites | | **Per-violation-type compliance test** | one test per legal requirement per subgroup, and again per subgroup comparison | Bouhoula et al. {[bouhoula2024_automated]} apply Holm–Bonferroni twice: to pairwise violation comparisons across popularity ranks, and to Fisher's exact tests comparing each CMP's subset against all websites |
-| **Per-regression-coefficient interaction term** | one test per coefficient, and the family is the model, not the paper | Nenadic et al. {[nenadic2026_overcoming]}: //"Because we estimate separate models for multiple disclosure obligations, conducting parallel hypothesis tests increases the risk of false positives"// — BH //"applied separately to the two families"// |+| **Per-regression-coefficient interaction term** | one test per coefficient, and the family is the model, not the paper | Nenadic et al. {[nenadic2026_swiss]}: //"Because we estimate separate models for multiple disclosure obligations, conducting parallel hypothesis tests increases the risk of false positives"// — BH //"applied separately to the two families"// |
 | **Per-site test** | one test per site. This is where //k// stops being tractable and the answer is a model, not a correction | none of the 269 papers with a correction tuple does this. Whether papers run per-site tests //without// correcting is not something this corpus was queried for | | **Per-site test** | one test per site. This is where //k// stops being tractable and the answer is a model, not a correction | none of the 269 papers with a correction tuple does this. Whether papers run per-site tests //without// correcting is not something this corpus was queried for |
  
Line 57: Line 57:
 The rule the corpus's better papers follow, stated as a rule: The rule the corpus's better papers follow, stated as a rule:
  
-  - **One family per confirmatory claim you will make in the abstract.** If the paper's claim is "consent rejection increases tracking", the family is the tests that support that claim — not every test in the paper. Nenadic et al. {[nenadic2026_overcoming]} correct //"separately to the two families of interaction terms"// because they make two claims.+  - **One family per confirmatory claim you will make in the abstract.** If the paper's claim is "consent rejection increases tracking", the family is the tests that support that claim — not every test in the paper. Nenadic et al. {[nenadic2026_swiss]} do this at the model level: because they //"estimate separate models for multiple disclosure obligations"//, they apply Benjamini–Hochberg //"separately to the two families of interaction terms (CH vs. EU and CH & EU vs. EU)"// — two contrast definitions of one research question, corrected apart rather than pooled.
   - **Then, per outcome variable.** Despres et al. {[despres2024_best]} fix α from //"a maximum of 25 hypotheses tested per outcome"//; Zimmeck et al. {[zimmeck2024_website]} correct across 231 comparisons drawn from one comparison set. Both state the grouping, which is what makes them checkable.   - **Then, per outcome variable.** Despres et al. {[despres2024_best]} fix α from //"a maximum of 25 hypotheses tested per outcome"//; Zimmeck et al. {[zimmeck2024_website]} correct across 231 comparisons drawn from one comparison set. Both state the grouping, which is what makes them checkable.
   - **A model is one family, not one test per coefficient across the paper.** If you fit six regressions, that is six families, each the size of its own coefficient set — not one family of all the coefficients you happen to report.   - **A model is one family, not one test per coefficient across the paper.** If you fit six regressions, that is six families, each the size of its own coefficient set — not one family of all the coefficients you happen to report.
Line 97: Line 97:
 **Nothing here has been overturned recently.** Multiplicity correction is one of the few methodological areas in web measurement where no 2020s development — LLM classification included — changed the answer. The procedures date from 1979, 1995 and 2001, and they are still the right ones. Where the field has moved is in //which// of them it reaches for, and slowly. Research-frontier work does exist — e-values and e-BH since 2022, selective inference — and **none of it has any presence in this corpus or any adopted standing**, so this page does not recommend it. **Nothing here has been overturned recently.** Multiplicity correction is one of the few methodological areas in web measurement where no 2020s development — LLM classification included — changed the answer. The procedures date from 1979, 1995 and 2001, and they are still the right ones. Where the field has moved is in //which// of them it reaches for, and slowly. Research-frontier work does exist — e-values and e-BH since 2022, selective inference — and **none of it has any presence in this corpus or any adopted standing**, so this page does not recommend it.
  
-What //has// moved is **reporting**, in an adjacent field. **CONSORT 2025** {[hopewell2025_consort]}the current reporting guideline for randomised trials, added multiplicity to what a paper must say: //"Any methods used to mitigate or account for multiplicity should be described. If no methods have been used to account for multiplicity (eg, not applicable, or not considered), then this should also be reported, particularly when a large number of analyses has been carried out."// That is exactly the standard this page argues for, it is dated 2025, and **none of the seven venues here asks for anything like it.**+What //has// moved is **reporting**, in an adjacent field, and by a smaller step than it first looksClinical trials have had multiplicity in their reporting guideline since CONSORT 2010, whose elaboration already discouraged multiple primary outcomes //"because of the problems of interpretation associated with multiplicity of analyses"// and whose limitations item already named //"multiplicity of analyses"//. What the **CONSORT 2025** elaboration {[hopewell2025_consort]} adds is the //explicit negative// obligation: //"Any methods used to mitigate or account for multiplicity should be described. If no methods have been used to account for multiplicity (eg, not applicable, or not considered), then this should also be reported, particularly when a large number of analyses has been carried out."//((The sentence is in the CONSORT 2025 **explanation and elaboration** document (BMJ 389:e081124, ''10.1136/bmj-2024-081124''), under the elaboration for item 21. It is **not** in the statement/checklist paper (''10.1136/bmj-2024-081123''), which contains no occurrence of "multiplicity" in its checklist text. Both fetched 2026-08-13.)) That is the standard the **Declining to correct** section of this page argues for, and **none of the seven venues here asks for anything like it** — nor does any of them ask for //k//, which CONSORT does not ask for either.
 </WRAP> </WRAP>
  
Line 106: Line 106:
   * **A compliance or enforcement claim** — "this CMP violates Article 7(3)" — is FWER. One false accusation is a problem in itself, and the audience includes people who will read exactly one row. Correct so that the probability of //any// false claim is bounded: Holm.   * **A compliance or enforcement claim** — "this CMP violates Article 7(3)" — is FWER. One false accusation is a problem in itself, and the audience includes people who will read exactly one row. Correct so that the probability of //any// false claim is bounded: Holm.
   * **A screen** — "which of these 400 third parties set an identifier before consent, so we can go look at them" — is FDR. You expect some of the shortlist to be wrong, you will follow up, and being unable to find anything is the real failure. Correct so that the //expected fraction// of wrong rows is bounded: Benjamini–Hochberg, or Benjamini–Yekutieli if you cannot argue positive dependence.   * **A screen** — "which of these 400 third parties set an identifier before consent, so we can go look at them" — is FDR. You expect some of the shortlist to be wrong, you will follow up, and being unable to find anything is the real failure. Correct so that the //expected fraction// of wrong rows is bounded: Benjamini–Hochberg, or Benjamini–Yekutieli if you cannot argue positive dependence.
-  * **A paper usually has both**, and the two families should be adjusted separately with the split stated. Nenadic et al. {[nenadic2026_overcoming]} do exactly this: BH //"applied separately to the two families of interaction terms"//. That is the pattern to copy — one correction across the whole paper is not more conservative, it is less meaningful.+  * **A paper usually has both**, and the two families should be adjusted separately with the split stated. Nenadic et al. {[nenadic2026_swiss]} do this for two contrast definitions of one question: BH //"applied separately to the two families of interaction terms"//. That is the pattern to copy — one correction across the whole paper is not more conservative, it is less meaningful.
  
 SciPy's own documentation states the trade-off in one line: FDR procedures //"tend to offer higher power than familywise error rate control procedures (e.g. Bonferroni correction)"//, and the ''by'' method //"is guaranteed to control the FDR even when the p-values are not from independent tests"//.((''scipy.stats.false_discovery_control'', SciPy 1.18.0 documentation, fetched 2026-08-13.)) SciPy's own documentation states the trade-off in one line: FDR procedures //"tend to offer higher power than familywise error rate control procedures (e.g. Bonferroni correction)"//, and the ''by'' method //"is guaranteed to control the FDR even when the p-values are not from independent tests"//.((''scipy.stats.false_discovery_control'', SciPy 1.18.0 documentation, fetched 2026-08-13.))
Line 376: Line 376:
 **No paper in this corpus engages with the idea.** The literal string //forking paths// matches two of the 5,869 full texts, and reading both shows they are packet-forwarding paths and symbolic-execution paths — see [[#What is missing entirely]]. The ASA statement on //p//-values {[wasserstein2016_asa]} matches **zero**. The standard mechanism that addresses it is committing to the analysis before you look — which is [[Statistics:Study preregistration]], and which **14 of these 1,025 hypothesis-test papers (1.4%) did**.((14, not the 6 in the ''statistics.kind'' table below. The schema field catches fewer than half the preregistrations in this corpus; the sibling page hand-classified all 62 full-text matches for ''pre-regist*'' and found 15 real study preregistrations, of which 14 also ran a hypothesis test. Use the hand count, not the field.)) It is not the only one: the multiverse analysis in [[#Open Questions]] addresses the same problem after the fact, and a held-out split of the crawl addresses part of it. **No paper in this corpus engages with the idea.** The literal string //forking paths// matches two of the 5,869 full texts, and reading both shows they are packet-forwarding paths and symbolic-execution paths — see [[#What is missing entirely]]. The ASA statement on //p//-values {[wasserstein2016_asa]} matches **zero**. The standard mechanism that addresses it is committing to the analysis before you look — which is [[Statistics:Study preregistration]], and which **14 of these 1,025 hypothesis-test papers (1.4%) did**.((14, not the 6 in the ''statistics.kind'' table below. The schema field catches fewer than half the preregistrations in this corpus; the sibling page hand-classified all 62 full-text matches for ''pre-regist*'' and found 15 real study preregistrations, of which 14 also ran a hypothesis test. Use the hand count, not the field.)) It is not the only one: the multiverse analysis in [[#Open Questions]] addresses the same problem after the fact, and a held-out split of the crawl addresses part of it.
  
-<wrap todo> +<WRAP todo> 
-The cheap version, if you will not preregister: **name your confirmatory tests in the paper**, correct within that family only, and label everything else exploratory with uncorrected //p//-values and no significance claims. Nenadic et al. {[nenadic2026_overcoming]} split their families; nobody in this corpus splits confirmatory from exploratory in a //crawl//+The cheap version, if you will not preregister: **name your confirmatory tests in the paper**, correct within that family only, and label everything else exploratory with uncorrected //p//-values and no significance claims. Nenadic et al. {[nenadic2026_swiss]} split their families; nobody in this corpus splits confirmatory from exploratory in a //crawl//
-</wrap>+</WRAP>
  
 ==== Declining to correct is a legitimate choice, if you say so ==== ==== Declining to correct is a legitimate choice, if you say so ====
Line 467: Line 467:
 ==== Almost nobody states the size of the family ==== ==== Almost nobody states the size of the family ====
  
-An adjusted //p// cannot be checked, reproduced or compared without //k//Of the 287 tuples that record real correction:+An adjusted //p// cannot be checked, reproduced or compared without //k//Two independent passes, agreeing to about point.
  
-^ What the paper states ^ Tuples ^ Share of 287 ^+**Pass 1, over the 287 tuples that record a real correction** — what landed in ''statistics[].method'', ''.detail'' or the evidence sentence: 
 + 
 +^ What the tuple states ^ Tuples ^ Share of 287 ^
 | a number of comparisons, tests or hypotheses in the family | **14** | **4.9%** | | a number of comparisons, tests or hypotheses in the family | **14** | **4.9%** |
 | an adjusted α or significance threshold | 57 | 19.9% | | an adjusted α or significance threshold | 57 | 19.9% |
 | any reported detail at all (''statistics[].detail'' non-null) | 161 | 56.1% | | any reported detail at all (''statistics[].detail'' non-null) | 161 | 56.1% |
 +
 +**Pass 2, over the whole text of the 265 papers that corrected.** A tuple only carries one sentence, so this pass searches the full paper for any count of comparisons, tests or hypotheses within ±1,500 characters of a procedure mention. Doing that naively is useless in this literature: the unguarded regex returned 41 hits, of which the majority were **χ² notation** — the digit in "χ² tests" sits immediately before the word — plus exponents and USENIX page furniture. With three guards the count is 20, and **reading all 20 leaves 15 genuine: 5.7% of the 265**. The five rejected, and every accepted match with its surrounding sentence, are printed in the report output on [[provenance:statistics:pvalue_corrections]].
 +
 +**So: 4.9% by tuple, 5.7% by paper. Roughly one corrected paper in twenty states the size of its family.**
  
 The fourteen that do state it are the ones a reader can check, and they read like this: The fourteen that do state it are the ones a reader can check, and they read like this:
Line 526: Line 532:
   * **Two buckets in the fold are not corrections** and are excluded from every "corrected" count: the three papers that declare they did not correct, and the two Greenhouse–Geisser sphericity corrections.   * **Two buckets in the fold are not corrections** and are excluded from every "corrected" count: the three papers that declare they did not correct, and the two Greenhouse–Geisser sphericity corrections.
   * **Sentinels are never answers.** ''statistics'' has no ''not-stated'' value — a paper that corrected and did not say so is simply absent, which is why every correction figure here is a **lower bound on the practice and an upper bound on the reporting**.   * **Sentinels are never answers.** ''statistics'' has no ''not-stated'' value — a paper that corrected and did not say so is simply absent, which is why every correction figure here is a **lower bound on the practice and an upper bound on the reporting**.
 +  * **The family-size figure is measured twice and hand-checked.** The tuple pass sees one sentence per correction; the full-text pass sees the whole paper but needs guards against χ² notation, exponents and page furniture, and then a reading of all 20 surviving hits. 4.9% and 5.7% are the two answers, five rejected hits are named with reasons, and all 20 are printed with their sentences on the provenance page. **Neither pass is a bound**: a count in a distant table caption is missed by both.
   * **A paper counts once**, never once per mention. 292 tuples across 269 papers.   * **A paper counts once**, never once per mention. 292 tuples across 269 papers.
   * **''statistics.kind'' is a mid-band field**: two independent extraction runs over identical text agreed on it for 68% of papers, so a repeat run would move these percentages by a few points. That caveat applies to every share on this page and does not apply to the folded procedure ranking, which is a ranking.   * **''statistics.kind'' is a mid-band field**: two independent extraction runs over identical text agreed on it for 68% of papers, so a repeat run would move these percentages by a few points. That caveat applies to every share on this page and does not apply to the folded procedure ranking, which is a ranking.
Line 548: Line 555:
 ===== Open Questions ===== ===== Open Questions =====
  
-<wrap todo>+<WRAP todo>
   * **No measurement paper in this corpus corrects a per-site family.** The design that most obviously creates thousands of hypotheses — one test per site — is never followed by a correction, so it is unknown whether the field considers that setting out of scope for testing, or simply does not test it. Both readings have consequences for how a reviewer should treat a per-site claim.   * **No measurement paper in this corpus corrects a per-site family.** The design that most obviously creates thousands of hypotheses — one test per site — is never followed by a correction, so it is unknown whether the field considers that setting out of scope for testing, or simply does not test it. Both readings have consequences for how a reviewer should treat a per-site claim.
   * **65 papers fit a multilevel model and none frames it as a multiplicity strategy** — at least, none surfaced while reading correction passages, and no one has checked all 65. Whether partial pooling is already doing the work of a correction in this literature without anyone saying so is answerable and unanswered.   * **65 papers fit a multilevel model and none frames it as a multiplicity strategy** — at least, none surfaced while reading correction passages, and no one has checked all 65. Whether partial pooling is already doing the work of a correction in this literature without anyone saying so is answerable and unanswered.
   * **Benjamini–Yekutieli is used by four papers**, and it is the procedure whose assumptions actually match site-level crawl data. Whether the BH results in this literature would survive BY is checkable on any paper that released its //p//-values — see [[:Artifacts]] — and nobody has done it.   * **Benjamini–Yekutieli is used by four papers**, and it is the procedure whose assumptions actually match site-level crawl data. Whether the BH results in this literature would survive BY is checkable on any paper that released its //p//-values — see [[:Artifacts]] — and nobody has done it.
   * **The debate is absent.** Zero papers cite the ASA statement, Rothman, Perneger, or the garden of forking paths. A short SoK on inference practice in web measurement would be citing an empty shelf, which is unusual and useful.   * **The debate is absent.** Zero papers cite the ASA statement, Rothman, Perneger, or the garden of forking paths. A short SoK on inference practice in web measurement would be citing an empty shelf, which is unusual and useful.
-  * **Nobody has asked these venues to require //k//.** Reporting the family size is a one-line checklist item that would make every correction in the literature checkable. CONSORT 2025 {[hopewell2025_consort]} now requires it of clinical trials; whether any of these seven venues would adopt the equivalent has not been proposed to them, and their review forms are not public, so the current state can only be read off the papers — where 84.5% of corrections come without a nearby //k//.+  * **Nobody has asked these venues to require //k//.** Reporting the family size is a one-line checklist item that would make every correction in the literature checkable, and **94.3% of corrected papers do not do it**. The nearest thing in any field is CONSORT's requirement that trials //describe the method// used — which is a weaker ask than stating //k//, and nobody has proposed even that much to a security or measurement venue. Their review forms are not public, so the current state can only be read off the papers.
   * **The pipeline-level multiplicity is unmeasured.** Nobody has taken a published crawl, enumerated the defensible alternatives at each pipeline decision (seed list, rank cut, exclusion rule, classifier threshold), re-run all of them, and reported the spread of the headline figure. That is a multiverse analysis, it is straightforwardly fundable, and it would say more about this literature's error rate than any correction.   * **The pipeline-level multiplicity is unmeasured.** Nobody has taken a published crawl, enumerated the defensible alternatives at each pipeline decision (seed list, rank cut, exclusion rule, classifier threshold), re-run all of them, and reported the spread of the headline figure. That is a multiverse analysis, it is straightforwardly fundable, and it would say more about this literature's error rate than any correction.
-</wrap>+</WRAP>
  
 ===== Related Pages ===== ===== Related Pages =====
statistics/pvalue_corrections.1786604990.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki