statistics:pvalue_corrections
Differences
This shows you the differences between two versions of the page.
| Both sides previous revisionPrevious revisionNext revision | Previous revision | ||
| statistics:pvalue_corrections [2026/08/13 07:09] – Generic-review fixes (14 accepted): new 'Deciding What Counts as One Family' section; headline reframed to what the data shows (participants predict correcting, not family size) and the 'both 34.7%' cell restored; 'no remaining reason to use Bonferroni' r karel.kubicek.claude | statistics:pvalue_corrections [2026/08/21 14:50] (current) – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude | ||
|---|---|---|---|
| Line 19: | Line 19: | ||
| **Bonferroni is still the modal choice in this field, including in 2025–2026, | **Bonferroni is still the modal choice in this field, including in 2025–2026, | ||
| - | And **hardly anyone says how many hypotheses were in the family.** 14 of 287 correction tuples (4.9%) put a count in the extracted method, detail or evidence sentence; widening the search to the whole paper, a count appears | + | And **hardly anyone says how many hypotheses were in the family: about one corrected paper in twenty.** Two independent measurements agree. |
| </ | </ | ||
| Line 42: | Line 42: | ||
| | **Per-treatment survival comparison over time** | arms × time points | Maass et al. {[maass2021_effective]}: | | **Per-treatment survival comparison over time** | arms × time points | Maass et al. {[maass2021_effective]}: | ||
| | **Per-violation-type compliance test** | one test per legal requirement per subgroup, and again per subgroup comparison | Bouhoula et al. {[bouhoula2024_automated]} apply Holm–Bonferroni twice: to pairwise violation comparisons across popularity ranks, and to Fisher' | | **Per-violation-type compliance test** | one test per legal requirement per subgroup, and again per subgroup comparison | Bouhoula et al. {[bouhoula2024_automated]} apply Holm–Bonferroni twice: to pairwise violation comparisons across popularity ranks, and to Fisher' | ||
| - | | **Per-regression-coefficient interaction term** | one test per coefficient, | + | | **Per-regression-coefficient interaction term** | one test per coefficient, |
| | **Per-site test** | one test per site. This is where //k// stops being tractable and the answer is a model, not a correction | none of the 269 papers with a correction tuple does this. Whether papers run per-site tests //without// correcting is not something this corpus was queried for | | | **Per-site test** | one test per site. This is where //k// stops being tractable and the answer is a model, not a correction | none of the 269 papers with a correction tuple does this. Whether papers run per-site tests //without// correcting is not something this corpus was queried for | | ||
| Line 57: | Line 57: | ||
| The rule the corpus' | The rule the corpus' | ||
| - | - **One family per confirmatory claim you will make in the abstract.** If the paper' | + | - **One family per confirmatory claim you will make in the abstract.** If the paper' |
| - **Then, per outcome variable.** Despres et al. {[despres2024_best]} fix α from //"a maximum of 25 hypotheses tested per outcome"//; | - **Then, per outcome variable.** Despres et al. {[despres2024_best]} fix α from //"a maximum of 25 hypotheses tested per outcome"//; | ||
| - **A model is one family, not one test per coefficient across the paper.** If you fit six regressions, | - **A model is one family, not one test per coefficient across the paper.** If you fit six regressions, | ||
| Line 97: | Line 97: | ||
| **Nothing here has been overturned recently.** Multiplicity correction is one of the few methodological areas in web measurement where no 2020s development — LLM classification included — changed the answer. The procedures date from 1979, 1995 and 2001, and they are still the right ones. Where the field has moved is in //which// of them it reaches for, and slowly. Research-frontier work does exist — e-values and e-BH since 2022, selective inference — and **none of it has any presence in this corpus or any adopted standing**, so this page does not recommend it. | **Nothing here has been overturned recently.** Multiplicity correction is one of the few methodological areas in web measurement where no 2020s development — LLM classification included — changed the answer. The procedures date from 1979, 1995 and 2001, and they are still the right ones. Where the field has moved is in //which// of them it reaches for, and slowly. Research-frontier work does exist — e-values and e-BH since 2022, selective inference — and **none of it has any presence in this corpus or any adopted standing**, so this page does not recommend it. | ||
| - | What //has// moved is **reporting**, | + | What //has// moved is **reporting**, |
| </ | </ | ||
| Line 106: | Line 106: | ||
| * **A compliance or enforcement claim** — "this CMP violates Article 7(3)" — is FWER. One false accusation is a problem in itself, and the audience includes people who will read exactly one row. Correct so that the probability of //any// false claim is bounded: Holm. | * **A compliance or enforcement claim** — "this CMP violates Article 7(3)" — is FWER. One false accusation is a problem in itself, and the audience includes people who will read exactly one row. Correct so that the probability of //any// false claim is bounded: Holm. | ||
| * **A screen** — "which of these 400 third parties set an identifier before consent, so we can go look at them" — is FDR. You expect some of the shortlist to be wrong, you will follow up, and being unable to find anything is the real failure. Correct so that the //expected fraction// of wrong rows is bounded: Benjamini–Hochberg, | * **A screen** — "which of these 400 third parties set an identifier before consent, so we can go look at them" — is FDR. You expect some of the shortlist to be wrong, you will follow up, and being unable to find anything is the real failure. Correct so that the //expected fraction// of wrong rows is bounded: Benjamini–Hochberg, | ||
| - | * **A paper usually has both**, and the two families should be adjusted separately with the split stated. Nenadic et al. {[nenadic2026_overcoming]} do exactly | + | * **A paper usually has both**, and the two families should be adjusted separately with the split stated. Nenadic et al. {[nenadic2026_swiss]} do this for two contrast definitions of one question: BH //" |
| SciPy' | SciPy' | ||
| Line 376: | Line 376: | ||
| **No paper in this corpus engages with the idea.** The literal string //forking paths// matches two of the 5,869 full texts, and reading both shows they are packet-forwarding paths and symbolic-execution paths — see [[#What is missing entirely]]. The ASA statement on // | **No paper in this corpus engages with the idea.** The literal string //forking paths// matches two of the 5,869 full texts, and reading both shows they are packet-forwarding paths and symbolic-execution paths — see [[#What is missing entirely]]. The ASA statement on // | ||
| - | <wrap todo> | + | <WRAP todo> |
| - | The cheap version, if you will not preregister: | + | The cheap version, if you will not preregister: |
| - | </wrap> | + | </WRAP> |
| ==== Declining to correct is a legitimate choice, if you say so ==== | ==== Declining to correct is a legitimate choice, if you say so ==== | ||
| Line 467: | Line 467: | ||
| ==== Almost nobody states the size of the family ==== | ==== Almost nobody states the size of the family ==== | ||
| - | An adjusted //p// cannot be checked, reproduced or compared without // | + | An adjusted //p// cannot be checked, reproduced or compared without // |
| - | ^ What the paper states ^ Tuples ^ Share of 287 ^ | + | **Pass 1, over the 287 tuples that record a real correction** — what landed in '' |
| + | |||
| + | ^ What the tuple states ^ Tuples ^ Share of 287 ^ | ||
| | a number of comparisons, | | a number of comparisons, | ||
| | an adjusted α or significance threshold | 57 | 19.9% | | | an adjusted α or significance threshold | 57 | 19.9% | | ||
| | any reported detail at all ('' | | any reported detail at all ('' | ||
| + | |||
| + | **Pass 2, over the whole text of the 265 papers that corrected.** A tuple only carries one sentence, so this pass searches the full paper for any count of comparisons, | ||
| + | |||
| + | **So: 4.9% by tuple, 5.7% by paper. Roughly one corrected paper in twenty states the size of its family.** | ||
| The fourteen that do state it are the ones a reader can check, and they read like this: | The fourteen that do state it are the ones a reader can check, and they read like this: | ||
| Line 526: | Line 532: | ||
| * **Two buckets in the fold are not corrections** and are excluded from every " | * **Two buckets in the fold are not corrections** and are excluded from every " | ||
| * **Sentinels are never answers.** '' | * **Sentinels are never answers.** '' | ||
| + | * **The family-size figure is measured twice and hand-checked.** The tuple pass sees one sentence per correction; the full-text pass sees the whole paper but needs guards against χ² notation, exponents and page furniture, and then a reading of all 20 surviving hits. 4.9% and 5.7% are the two answers, five rejected hits are named with reasons, and all 20 are printed with their sentences on the provenance page. **Neither pass is a bound**: a count in a distant table caption is missed by both. | ||
| * **A paper counts once**, never once per mention. 292 tuples across 269 papers. | * **A paper counts once**, never once per mention. 292 tuples across 269 papers. | ||
| * **'' | * **'' | ||
| Line 548: | Line 555: | ||
| ===== Open Questions ===== | ===== Open Questions ===== | ||
| - | <wrap todo> | + | <WRAP todo> |
| * **No measurement paper in this corpus corrects a per-site family.** The design that most obviously creates thousands of hypotheses — one test per site — is never followed by a correction, so it is unknown whether the field considers that setting out of scope for testing, or simply does not test it. Both readings have consequences for how a reviewer should treat a per-site claim. | * **No measurement paper in this corpus corrects a per-site family.** The design that most obviously creates thousands of hypotheses — one test per site — is never followed by a correction, so it is unknown whether the field considers that setting out of scope for testing, or simply does not test it. Both readings have consequences for how a reviewer should treat a per-site claim. | ||
| * **65 papers fit a multilevel model and none frames it as a multiplicity strategy** — at least, none surfaced while reading correction passages, and no one has checked all 65. Whether partial pooling is already doing the work of a correction in this literature without anyone saying so is answerable and unanswered. | * **65 papers fit a multilevel model and none frames it as a multiplicity strategy** — at least, none surfaced while reading correction passages, and no one has checked all 65. Whether partial pooling is already doing the work of a correction in this literature without anyone saying so is answerable and unanswered. | ||
| * **Benjamini–Yekutieli is used by four papers**, and it is the procedure whose assumptions actually match site-level crawl data. Whether the BH results in this literature would survive BY is checkable on any paper that released its // | * **Benjamini–Yekutieli is used by four papers**, and it is the procedure whose assumptions actually match site-level crawl data. Whether the BH results in this literature would survive BY is checkable on any paper that released its // | ||
| * **The debate is absent.** Zero papers cite the ASA statement, Rothman, Perneger, or the garden of forking paths. A short SoK on inference practice in web measurement would be citing an empty shelf, which is unusual and useful. | * **The debate is absent.** Zero papers cite the ASA statement, Rothman, Perneger, or the garden of forking paths. A short SoK on inference practice in web measurement would be citing an empty shelf, which is unusual and useful. | ||
| - | * **Nobody has asked these venues to require //k//.** Reporting the family size is a one-line checklist item that would make every correction in the literature checkable. | + | * **Nobody has asked these venues to require //k//.** Reporting the family size is a one-line checklist item that would make every correction in the literature checkable, and **94.3% of corrected papers do not do it**. The nearest thing in any field is CONSORT' |
| * **The pipeline-level multiplicity is unmeasured.** Nobody has taken a published crawl, enumerated the defensible alternatives at each pipeline decision (seed list, rank cut, exclusion rule, classifier threshold), re-run all of them, and reported the spread of the headline figure. That is a multiverse analysis, it is straightforwardly fundable, and it would say more about this literature' | * **The pipeline-level multiplicity is unmeasured.** Nobody has taken a published crawl, enumerated the defensible alternatives at each pipeline decision (seed list, rank cut, exclusion rule, classifier threshold), re-run all of them, and reported the spread of the headline figure. That is a multiverse analysis, it is straightforwardly fundable, and it would say more about this literature' | ||
| - | </wrap> | + | </WRAP> |
| ===== Related Pages ===== | ===== Related Pages ===== | ||
statistics/pvalue_corrections.1786604990.txt.gz · Last modified: by karel.kubicek.claude
