Context Theory Get your growth audit

Answer

Should verification be performed by a separate agent?

Yes, if it sees the artefact and the requirement without the reasoning that produced them. Otherwise it agrees.

Yes, when it sees only the output and the requirement. A reviewer given the full working history re-derives the same conclusions and agrees, so separation without independence buys a second confident opinion and nothing else.

The property doing the work is not that a different run performs the check; it is that the checking run does not hold the reasoning that produced the artefact. Given the full history, a reviewer follows the same path and reaches the same conclusion, which is why an in-session review returns approval so reliably. Given only the output and the requirement, it has to evaluate the thing itself, and that is a different operation with a different failure rate.

This sets the design directly. The verification pass receives the requirement as originally stated, the artefact, and nothing else. Not the conversation, not the plan, not the explanation of why particular choices were made. Withholding those feels like withholding useful context and is precisely the point: the explanation is the thing most likely to persuade a reviewer that a wrong output is right.

The strongest use is checking against a stated rubric rather than asking whether the work is good. Does the output satisfy each stated requirement, item by item, with evidence. Are there requirements it does not address. Does it contain claims not supported by anything supplied. These are discrimination tasks and are performed considerably better than an open judgement of quality.

There are limits worth stating. A verification pass shares the tendencies of the run that produced the work, so an error arising from a common assumption about the domain will be reproduced rather than caught. It also cannot detect what was never attempted unless coverage is part of what it checks, which means the requirement supplied to it has to include the scope and not only the standard.

Where a deterministic check exists, it beats both. A test, a schema validation, a query returning nothing, a comparison against a known-good output: each of these produces the same verdict every time and cannot be persuaded. The order of preference is deterministic check, then independent model pass, then a review by the run that did the work, and the third is worth very little.

Finally, the verification pass should be able to fail. If its output is a summary that always concludes the work is acceptable with minor observations, it is producing the shape of a review rather than a verdict. Requiring an explicit pass or fail against each requirement, and treating a failure as an outcome rather than as an obstacle, is what keeps it a check.

A reviewer who watched the work being done is not a second opinion, it is the first opinion with a different name.

Siddharth Sharma, Context Theory

Related questions

Does using a different model help?

Somewhat, for errors that stem from one system's particular habits, and not for errors that stem from a shared misreading of the domain. The larger effect comes from withholding the reasoning rather than from changing the reviewer, which is worth knowing because it is free and the other is not.

Should the verification pass be able to fix what it finds?

No, and separating the two is what keeps the finding trustworthy. A reviewer that repairs its own findings produces an artefact nobody has independently assessed, and the incentive to resolve rather than report is strong. Report, then decide, then fix in a separate step.

METHOD

Every figure below carries its source and the date it was verified. Nothing on this page is asserted.

The numbers on this page.

Datapoints
What Value Specific to
Close rate — response under 5 minutes vs over 24 hours32% vs 12%Category-wide
Sub-15-minute compliance — automated routing vs manual only62.5% vs 39.1%Category-wide

Optifai speed-to-lead benchmark · n=939 companies · Q2 2025–Q1 2026 · verified

2026 speed-to-lead benchmark · verified

What is specific to this page.

Evidence
Kind Claim Check it against
WorkflowIndependence rather than separation makes a verification pass informative, because a reviewer holding the originating reasoning follows the same path to the same conclusion, which is why in-session review returns approval at a high rate.Comparing verdicts from a reviewer given the full session against one given only the artefact and the requirement.
SoftwareChecking an artefact item by item against a stated rubric is a discrimination task and is performed markedly better than an open assessment of quality, which is why the requirement must be supplied in enumerable form.Comparing findings from a rubric-based check against an open request to evaluate the same output.
ConstraintA verification pass cannot detect work never attempted unless coverage is part of what it checks, so the material supplied to it must state the scope as well as the standard.Submitting an output covering part of the required scope and observing whether the check reports the omission.
ResponseA deterministic check produces the same verdict on every run and cannot be persuaded, which places it above an independent model pass, which in turn is above a review by the run that produced the work.Running all three against an output containing a known defect and comparing which detect it consistently.

Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.

Start with the measurement.

Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.

Get your growth audit

$497 · delivered in 5 business days · credited against month one