User Tools

Site Tools


statistics:pvalue_corrections

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
statistics:pvalue_corrections [2026/08/13 07:25] – One more artefact found in the family-size hand list by reading the source: the PETS 2022 multi-region hit was running-header furniture, not a count. Figure moves from 16/265 (6.0%) to 15/265 (5.7%), complement to 94.3%. Authored by Claude karel.kubicek.claudestatistics:pvalue_corrections [2026/08/21 14:50] (current) – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude
Line 376: Line 376:
 **No paper in this corpus engages with the idea.** The literal string //forking paths// matches two of the 5,869 full texts, and reading both shows they are packet-forwarding paths and symbolic-execution paths — see [[#What is missing entirely]]. The ASA statement on //p//-values {[wasserstein2016_asa]} matches **zero**. The standard mechanism that addresses it is committing to the analysis before you look — which is [[Statistics:Study preregistration]], and which **14 of these 1,025 hypothesis-test papers (1.4%) did**.((14, not the 6 in the ''statistics.kind'' table below. The schema field catches fewer than half the preregistrations in this corpus; the sibling page hand-classified all 62 full-text matches for ''pre-regist*'' and found 15 real study preregistrations, of which 14 also ran a hypothesis test. Use the hand count, not the field.)) It is not the only one: the multiverse analysis in [[#Open Questions]] addresses the same problem after the fact, and a held-out split of the crawl addresses part of it. **No paper in this corpus engages with the idea.** The literal string //forking paths// matches two of the 5,869 full texts, and reading both shows they are packet-forwarding paths and symbolic-execution paths — see [[#What is missing entirely]]. The ASA statement on //p//-values {[wasserstein2016_asa]} matches **zero**. The standard mechanism that addresses it is committing to the analysis before you look — which is [[Statistics:Study preregistration]], and which **14 of these 1,025 hypothesis-test papers (1.4%) did**.((14, not the 6 in the ''statistics.kind'' table below. The schema field catches fewer than half the preregistrations in this corpus; the sibling page hand-classified all 62 full-text matches for ''pre-regist*'' and found 15 real study preregistrations, of which 14 also ran a hypothesis test. Use the hand count, not the field.)) It is not the only one: the multiverse analysis in [[#Open Questions]] addresses the same problem after the fact, and a held-out split of the crawl addresses part of it.
  
-<wrap todo>+<WRAP todo>
 The cheap version, if you will not preregister: **name your confirmatory tests in the paper**, correct within that family only, and label everything else exploratory with uncorrected //p//-values and no significance claims. Nenadic et al. {[nenadic2026_swiss]} split their families; nobody in this corpus splits confirmatory from exploratory in a //crawl//. The cheap version, if you will not preregister: **name your confirmatory tests in the paper**, correct within that family only, and label everything else exploratory with uncorrected //p//-values and no significance claims. Nenadic et al. {[nenadic2026_swiss]} split their families; nobody in this corpus splits confirmatory from exploratory in a //crawl//.
-</wrap>+</WRAP>
  
 ==== Declining to correct is a legitimate choice, if you say so ==== ==== Declining to correct is a legitimate choice, if you say so ====
Line 555: Line 555:
 ===== Open Questions ===== ===== Open Questions =====
  
-<wrap todo>+<WRAP todo>
   * **No measurement paper in this corpus corrects a per-site family.** The design that most obviously creates thousands of hypotheses — one test per site — is never followed by a correction, so it is unknown whether the field considers that setting out of scope for testing, or simply does not test it. Both readings have consequences for how a reviewer should treat a per-site claim.   * **No measurement paper in this corpus corrects a per-site family.** The design that most obviously creates thousands of hypotheses — one test per site — is never followed by a correction, so it is unknown whether the field considers that setting out of scope for testing, or simply does not test it. Both readings have consequences for how a reviewer should treat a per-site claim.
   * **65 papers fit a multilevel model and none frames it as a multiplicity strategy** — at least, none surfaced while reading correction passages, and no one has checked all 65. Whether partial pooling is already doing the work of a correction in this literature without anyone saying so is answerable and unanswered.   * **65 papers fit a multilevel model and none frames it as a multiplicity strategy** — at least, none surfaced while reading correction passages, and no one has checked all 65. Whether partial pooling is already doing the work of a correction in this literature without anyone saying so is answerable and unanswered.
Line 562: Line 562:
   * **Nobody has asked these venues to require //k//.** Reporting the family size is a one-line checklist item that would make every correction in the literature checkable, and **94.3% of corrected papers do not do it**. The nearest thing in any field is CONSORT's requirement that trials //describe the method// used — which is a weaker ask than stating //k//, and nobody has proposed even that much to a security or measurement venue. Their review forms are not public, so the current state can only be read off the papers.   * **Nobody has asked these venues to require //k//.** Reporting the family size is a one-line checklist item that would make every correction in the literature checkable, and **94.3% of corrected papers do not do it**. The nearest thing in any field is CONSORT's requirement that trials //describe the method// used — which is a weaker ask than stating //k//, and nobody has proposed even that much to a security or measurement venue. Their review forms are not public, so the current state can only be read off the papers.
   * **The pipeline-level multiplicity is unmeasured.** Nobody has taken a published crawl, enumerated the defensible alternatives at each pipeline decision (seed list, rank cut, exclusion rule, classifier threshold), re-run all of them, and reported the spread of the headline figure. That is a multiverse analysis, it is straightforwardly fundable, and it would say more about this literature's error rate than any correction.   * **The pipeline-level multiplicity is unmeasured.** Nobody has taken a published crawl, enumerated the defensible alternatives at each pipeline decision (seed list, rank cut, exclusion rule, classifier threshold), re-run all of them, and reported the spread of the headline figure. That is a multiverse analysis, it is straightforwardly fundable, and it would say more about this literature's error rate than any correction.
-</wrap>+</WRAP>
  
 ===== Related Pages ===== ===== Related Pages =====
statistics/pvalue_corrections.1786605930.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki