Methods · Oct 1, 2026 · 5 min read
Agreement is a measurement, not a feeling
Teams often say their reviewers 'mostly agree'. That sentence hides a lot. If two raters pick between two options at random, they will still agree about half the time.
Correct for chance
Measures such as Cohen's kappa adjust raw agreement for the agreement you would expect by luck. A high raw score with a low adjusted score means the task is easier to guess than it looks.
Low agreement usually means a vague rubric
When experts disagree, the instinct is to blame the experts. More often the criteria are open to interpretation. Words like 'good' and 'appropriate' invite different readings.
What to do about it
Track agreement per criterion, not just overall. Find the criteria where reviewers split, rewrite them until they can be applied the same way, and measure again. Agreement is something you build, then verify.
Need expert judgment on your model's output?
Try SoReliable