Answer
How do you reconcile conflicting results from agents?
Go back to what each one read. A disagreement between two summaries cannot be settled by comparing the summaries.
Return to the evidence each one used. A conflict between two summaries is not resolvable from the summaries, and asking a third run to decide produces a confident answer with no additional information behind it.
The instinct is to weigh the two answers: which is more detailed, more confident, more plausible. None of these correlates with correctness, and all of them are properties of how the answer was written rather than of what was found. A disagreement is a signal that at least one run saw something the other did not, or interpreted the same thing differently, and both possibilities are only distinguishable at the level of evidence.
So the first move is to look at what each read. This is only possible if the results carry references, which is the practical argument for requiring them: a conflict between two referenced answers is usually resolved in a minute by opening both sources, and a conflict between two unreferenced answers can only be resolved by redoing the work. The requirement pays for itself the first time this happens.
The most common resolution is that the two looked at different instances of the same thing. Two configurations, two environments, two versions of a document, two records for the same customer. Neither run was wrong; the question had more than one answer and neither knew that. This resolution is valuable beyond the immediate conflict, because it has surfaced an ambiguity in the system that was previously invisible.
The second most common is that one run inferred where the other observed. A conclusion drawn from what a value implies conflicts with a conclusion drawn from reading the value, and the disagreement disappears the moment the two are labelled. This is the argument for results distinguishing what was seen from what was concluded, and the argument arrives most forcefully during a conflict.
Asking a third run to adjudicate is the tempting move and it is worth being clear about what it does. The arbiter has no access to the evidence unless it is given, so it is choosing between two pieces of writing, and it will choose the more confident and better organised one. Where the arbiter is given the underlying evidence, it is not arbitrating at all — it is doing the work a third time, which is a legitimate thing to do and should be described accurately.
Finally, a conflict is worth recording rather than merely resolving. It marks a place where the answer was not obvious, and the same question will be asked again by a later session. Writing down what the conflict was, what the evidence showed, and which reading is correct converts an hour of investigation into a line, and it is the kind of finding that exists nowhere else once the session ends.
Two disagreeing summaries and a third opinion give you three opinions, which is one fewer fact than you had.
Siddharth Sharma, Context Theory
Related questions
What if both results are wrong?
It happens, and a conflict is a good moment to notice it, because the disagreement invites the check that agreement never prompts. Two runs agreeing is weaker evidence than it feels, particularly where both read the same material: agreement can mean the answer is right or that both inherited the same error.
Should agents be told to flag uncertainty to prevent conflicts?
Flagging uncertainty is useful and does not prevent this, because most conflicts arise where each run was legitimately confident about the part it saw. The preventive measure is per-claim evidence, which makes a conflict cheap to resolve rather than making it less likely.
METHOD
Every figure below carries its source and the date it was verified. Nothing on this page is asserted.
The numbers on this page.
| What | Value | Specific to |
|---|---|---|
| Close rate — response under 5 minutes vs over 24 hours | 32% vs 12% | Category-wide |
| Sub-15-minute compliance — automated routing vs manual only | 62.5% vs 39.1% | Category-wide |
Optifai speed-to-lead benchmark · n=939 companies · Q2 2025–Q1 2026 · verified
2026 speed-to-lead benchmark · verified
What is specific to this page.
| Kind | Claim | Check it against |
|---|---|---|
| Workflow | Detail, confidence and plausibility are properties of how an answer was written rather than of what was found, so weighing two conflicting summaries against each other cannot establish which is correct. | Comparing the confidence and length of conflicting results against which one the underlying evidence supported. |
| Software | The most common resolution is that the runs examined different instances of the same thing — two configurations, environments, versions or records — which means neither was wrong and an ambiguity in the system has been surfaced. | Comparing the specific sources each run cited for the conflicting claim. |
| Response | A third run given only the two answers is choosing between pieces of writing and will favour the more confident and better organised one, while a third run given the evidence is repeating the work rather than arbitrating. | Supplying an arbiter with two conflicting summaries and observing whether its choice tracks presentation or evidence. |
| Constraint | Agreement between two runs reading the same material is weaker evidence than it appears, because it is equally consistent with both having inherited the same error from the source. | Checking whether two agreeing runs relied on the same underlying document. |
Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.
Start with the measurement.
Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.
$497 · delivered in 5 business days · credited against month one