Evaluation for judgement work
Ranking candidates, applying a commission plan, deciding what counts — none of these have a clean ground truth, so the usual benchmark loop does not apply.
A benchmark needs an answer key. Judgement work does not have one: two good recruiters produce different shortlists from the same mandate and both can be right, and a commission dispute is usually about which deals counted rather than about arithmetic.
So the question is what to measure instead. Agreement between independent reviewers is one candidate. Whether a decision survives being explained to the person it affects is another, and it is closer to what the work actually demands.
We do not have a good answer to this yet. It is the question that most often decides whether a feature ships.