User Tools

Site Tools


statistics:interrater_agreement

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
statistics:interrater_agreement [2026/08/21 08:33] – [Open Questions] karelkubicekstatistics:interrater_agreement [2026/08/21 14:50] (current) – Boxes: <wrap> renders a span, use uppercase <WRAP>; one box per list, not per bullet karel.kubicek.claude
Line 108: Line 108:
 </code> </code>
  
-<wrap todo>+<WRAP todo>
 **With more than two coders, use Fleiss' kappa (equal number of coders per item, nominal labels) or Krippendorff's alpha (unequal coders, missing data, or ordinal/interval labels).** If you average pairwise Cohen's kappas anyway, say so and give the range, not the mean alone — the mean of three pairwise kappas is not "the" kappa. **With more than two coders, use Fleiss' kappa (equal number of coders per item, nominal labels) or Krippendorff's alpha (unequal coders, missing data, or ordinal/interval labels).** If you average pairwise Cohen's kappas anyway, say so and give the range, not the mean alone — the mean of three pairwise kappas is not "the" kappa.
-</wrap>+</WRAP>
  
 Krippendorff's alpha is the safest default for the messy case, which is why the more careful papers here reach for it: Cory et al. {[cory2026_wordlevel]} compute it over //"a team of six domain experts from two universities"// on 200 privacy policies (α = 0.669); Xian et al. {[xian2025_layered]} use the unitising variant Cu-α for span-level policy annotation, reporting 0.95 for CCPA policies and 0.89 for the rest. Krippendorff's alpha is the safest default for the messy case, which is why the more careful papers here reach for it: Cory et al. {[cory2026_wordlevel]} compute it over //"a team of six domain experts from two universities"// on 200 privacy policies (α = 0.669); Xian et al. {[xian2025_layered]} use the unitising variant Cu-α for span-level policy annotation, reporting 0.95 for CCPA policies and 0.89 for the rest.
Line 256: Line 256:
 Read the first table with a real task in mind. If 3% of your sampled cookies are the class you care about and your two coders disagree on 3% of items, your kappa is about **0.65** — a number a reviewer will read as mediocre — and the same two coders on a balanced sample score **0.94**. Nothing about their reliability changed. Read the first table with a real task in mind. If 3% of your sampled cookies are the class you care about and your two coders disagree on 3% of items, your kappa is about **0.65** — a number a reviewer will read as mediocre — and the same two coders on a balanced sample score **0.94**. Nothing about their reliability changed.
  
-<wrap todo>+<WRAP todo>
 **On a lopsided task, report three things, not one:** the coefficient, the **observed agreement**, and the **marginal distribution** (how many items each coder put in each class). Those three together let a reader diagnose the paradox; any one alone does not. If the class is very rare, report **Gwet's AC1** {[gwet2008_computing]} alongside kappa — it is designed for this case, it is in ''irrCAC'' on CRAN, and five papers in this corpus already do it. **On a lopsided task, report three things, not one:** the coefficient, the **observed agreement**, and the **marginal distribution** (how many items each coder put in each class). Those three together let a reader diagnose the paradox; any one alone does not. If the class is very rare, report **Gwet's AC1** {[gwet2008_computing]} alongside kappa — it is designed for this case, it is in ''irrCAC'' on CRAN, and five papers in this corpus already do it.
  
 The corpus's model of this is Khatun et al. {[khatun2026_disclosure]}, who give both numbers for the same 40-app subsample: Cohen's κ = 0.78 **and** 90% observed agreement. Chen et al. {[chen2025_semantics]} do the same job at scale on the task this site cares about most — cookie purpose labels, three independent author-coders, 2,300 cookies, Fleiss' κ = 0.978. The corpus's model of this is Khatun et al. {[khatun2026_disclosure]}, who give both numbers for the same 40-app subsample: Cohen's κ = 0.78 **and** 90% observed agreement. Chen et al. {[chen2025_semantics]} do the same job at scale on the task this site cares about most — cookie purpose labels, three independent author-coders, 2,300 cookies, Fleiss' κ = 0.978.
-</wrap>+</WRAP>
  
 ==== Landis–Koch bands are not a result ==== ==== Landis–Koch bands are not a result ====
statistics/interrater_agreement.1787301225.txt.gz · Last modified: by karelkubicek

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki