Blog & Research
Notes on evaluating expert work in AI systems: methods, case studies and open questions.
Software · Oct 6, 2026 · 4 min read
Passing tests is not the same as working code
Automated checks only prove what they run. The bugs that matter live in what nobody tested.
Read the articleMethods · Oct 1, 2026
Agreement is a measurement, not a feeling
Two reviewers who agree might just be guessing alike. Here is how to tell.
Robotics · Sep 24, 2026
Reviewing robot plans before they run
A plan that reads well on screen can still crush a glass vial or cross a hot plate.
Clinical · Sep 17, 2026
Fluent is not correct: reviewing AI in clinical settings
The hardest errors to catch are omissions. A smooth answer can leave out the one thing that matters.
Methods · Sep 9, 2026
The anatomy of a good rubric
Atomic, binary, observable. How to turn a quality bar into criteria people can apply the same way.
Company · Sep 2, 2026
Plan, assess, train, test: how we qualify specialists
A look at the process behind SoReliable projects, and why every step comes before live work.
Our research agenda
Questions we are working on and writing about next.
- →How much does expert review add beyond automated judges?
- →Which rubric designs give the highest agreement between reviewers?
- →How should physical safety be scored for robot plans?
- →What do omissions in clinical AI look like, and how can they be measured?
Ready to try SoReliable?
Tell us what you are building. Any team, any project, any size. We find the specialists, train them on your standard, and evaluate the work against what your project actually needs.