Context Theory Get your growth audit

Answer

Should you use a model to grade a model?

For comparisons against a stated criterion, usefully. For an open judgement of quality, it mostly measures how well the output is written.

For checking an output against explicit criteria, yes, and it is often the only thing that scales. For open quality judgement, it largely measures fluency and length, which are the properties least related to whether the output is correct.

The distinction that decides this is between discrimination and assessment. Discrimination is a comparison against something stated: does this output satisfy this requirement, does this passage support this claim, does this answer contain the value from that record. Assessment is an open judgement: is this good, is this thorough, is this well reasoned. The first is performed usefully and reproducibly. The second correlates with surface properties, and the surface is what the system producing the output is best at.

That gives a design rule. Never ask for a score without a rubric, and make the rubric a list of conditions rather than a set of qualities. Each item should be answerable with a yes, a no and a pointer to the evidence. A grader that returns per-item verdicts with references can be spot-checked; one that returns a number out of ten cannot be, and the number will drift with the phrasing of what it is grading.

Independence is the second requirement and it is the one usually missing. A grader that sees the reasoning which produced the output re-derives the same conclusions and agrees, which is why in-conversation self-grading returns approval at a high rate. Supplying only the artefact and the criteria is what makes the verdict informative, and withholding the reasoning is deliberate rather than an oversight.

Known biases are worth designing around because they are consistent. Longer answers score higher than shorter ones of equal content. Answers matching the grader's own habitual structure score higher. Position affects comparative judgements, so the same pair evaluated in both orders is worth doing where the comparison matters. None of these disqualifies the method; each of them means an unstructured score is measuring something other than what it claims.

The strongest use in practice is as a filter rather than a verdict. Run the grader across everything, use it to surface the cases that fail a stated criterion, and have a person look at those. This uses discrimination where it is strong and reserves judgement for a small set, and it produces a system whose weakest link is a human reading twenty items rather than a number nobody can check.

Finally, validate the grader itself before relying on it. Give it a set of outputs whose correctness you already know, including some that are wrong in the specific ways you care about, and see whether it catches them. A grader that has never been tested against known-bad cases is an unmeasured component being used to measure other things, which is not an improvement over having no measurement.

Ask a model whether this is good and it will tell you how well it is written, which is a real answer to a question you did not ask.

Siddharth Sharma, Context Theory

Related questions

Does using a different model as grader help?

It removes some shared habits and not the structural biases, which appear across systems: length, fluency and structural familiarity all affect judgement regardless of which model is grading. The larger improvement comes from withholding the reasoning and supplying explicit criteria, both of which are free.

Is human grading better?

More reliable per item and unavailable at volume, which is the whole reason this question exists. The productive arrangement uses human grading to calibrate and audit the model grader rather than to replace it: a small sample graded by a person, compared against the model's verdicts, tells you what the automated number is worth.

METHOD

Every figure below carries its source and the date it was verified. Nothing on this page is asserted.

The numbers on this page.

Datapoints
What Value Specific to
Close rate — response under 5 minutes vs over 24 hours32% vs 12%Category-wide
Sub-15-minute compliance — automated routing vs manual only62.5% vs 39.1%Category-wide

Optifai speed-to-lead benchmark · n=939 companies · Q2 2025–Q1 2026 · verified

2026 speed-to-lead benchmark · verified

What is specific to this page.

Evidence
Kind Claim Check it against
WorkflowModel grading performs discrimination against stated criteria reproducibly and performs open quality assessment as a correlate of surface properties, which are precisely what the generating system optimises.Grading the same content presented well and presented poorly under an open quality prompt and under an explicit rubric.
SoftwareA grader supplied with the reasoning that produced an output re-derives the same conclusions and agrees, so withholding that reasoning is what makes the verdict informative rather than an oversight in the design.Comparing grader verdicts with and without the originating conversation supplied.
ResponseConsistent biases affect model grading — longer answers over shorter ones of equal content, familiar structure over unfamiliar, and position in comparative judgements — so an unstructured score measures something other than what it claims.Grading an identical pair in both presentation orders and comparing the outcomes.
ConstraintA grader that has not been tested against outputs known to be wrong in the ways that matter is an unmeasured component being used to measure others, which does not improve on having no measurement.Running the grader over a set of known-bad outputs and counting how many it flags.

Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.

Start with the measurement.

Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.

Get your growth audit

$497 · delivered in 5 business days · credited against month one