All articles

Methods · Oct 1, 2026 · 5 min read

Agreement is a measurement, not a feeling

Teams often say their reviewers 'mostly agree'. That sentence hides a lot. If two raters pick between two options at random, they will still agree about half the time.

Correct for chance

Measures such as Cohen's kappa adjust raw agreement for the agreement you would expect by luck. A high raw score with a low adjusted score means the task is easier to guess than it looks.

Low agreement usually means a vague rubric

When experts disagree, the instinct is to blame the experts. More often the criteria are open to interpretation. Words like 'good' and 'appropriate' invite different readings.

What to do about it

Track agreement per criterion, not just overall. Find the criteria where reviewers split, rewrite them until they can be applied the same way, and measure again. Agreement is something you build, then verify.

Need expert judgment on your model's output?

Try SoReliable

KEEP READING

Ready to try SoReliable?

Tell us what you are building. Any team, any project, any size. We find the specialists, train them on your standard, and evaluate the work against what your project actually needs.

You'll hear back from our solutions team, not an autoresponder.