Research Thread 02

Evaluation for judgement work

Ranking candidates, applying a commission plan, deciding what counts — none of these have a clean ground truth, so the usual benchmark loop does not apply.

A benchmark needs an answer key. Judgement work does not have one: two good recruiters produce different shortlists from the same mandate and both can be right, and a commission dispute is usually about which deals counted rather than about arithmetic.

So the question is what to measure instead. Agreement between independent reviewers is one candidate. Whether a decision survives being explained to the person it affects is another, and it is closer to what the work actually demands.

We do not have a good answer to this yet. It is the question that most often decides whether a feature ships.

01 Provenance as a product constraint 03 Small models inside a boundary