Context Theory Get your growth audit

Answer

How do you use AI with documents you receive?

Extract the fields you need into a record, and keep the document. The extraction is checkable; a summary of it is not.

Extract the specific fields you need into a record and keep the original alongside. Extraction is checkable because the value sits next to its source; a summary replaces the document with an account nobody can verify at a glance.

Businesses receive documents constantly — invoices, quotes, orders, specifications, statements, certificates — and the work they create is the same each time: read it, take out the handful of facts that matter, put them somewhere. That is an extraction task with a defined output, which is the shape these systems handle most reliably and the shape that is easiest to check.

Defining the fields is most of the work and it is worth doing carefully. Supplier, date, reference, total, tax, due date, line items: the specific list depends on what you do with the document afterwards. A field that nothing downstream uses should not be extracted, and a field that something downstream requires must have a defined behaviour when it is absent from the document, which happens more often than any specification anticipates.

The original must be kept and linked. This is what makes the extraction verifiable — a person or an auditor can compare the field against the source in seconds — and it is also what makes the arrangement recoverable when the extraction turns out to have been wrong for a period. A system that extracts and discards has converted a document into an assertion.

Validation belongs between extraction and use. Does the total match the sum of the lines. Is the date plausible. Does the supplier reference resolve to a supplier you have. Is the currency one you deal in. Each of these catches a class of extraction error mechanically, and each is a few lines of code. Where a validation fails, the item should go to a person rather than proceeding with a corrected guess.

Scanned and photographed documents are the case that behaves differently and should be handled explicitly. Quality varies, text may be absent entirely, and confidence in a value read from a poor image is genuinely lower. Treating these as a separate stream with mandatory review is more honest than a single path whose accuracy silently depends on how the document arrived.

Summarising a document is a different operation and a weaker one for this purpose. A summary is not checkable at a glance, cannot be validated, and cannot be used by anything downstream that needs a value. Summaries are useful for orientation — what is this document about, does it need attention — and should not be the thing that enters your records.

Extract fields, not meaning, because a field can be wrong in a way you can see.

Siddharth Sharma, Context Theory

Related questions

What accuracy is achievable on document extraction?

It depends far more on document consistency than on anything else. A supplier who sends the same layout every month is close to solved; a mixed stream of formats from many senders is not, and the difference is large. The useful measurement is per sender rather than overall, because it points at where the manual handling should stay.

Should extracted values be corrected in place or flagged?

Flagged and corrected by a person, with the correction recorded. Correcting in place silently means nobody learns which senders or fields are causing trouble, and that information is what tells you where to fix the process — often by asking a supplier for a different format, which solves it permanently.

METHOD

Every figure below carries its source and the date it was verified. Nothing on this page is asserted.

The numbers on this page.

Datapoints
What Value Specific to
Close rate — response under 5 minutes vs over 24 hours32% vs 12%Category-wide
Sub-15-minute compliance — automated routing vs manual only62.5% vs 39.1%Category-wide

Optifai speed-to-lead benchmark · n=939 companies · Q2 2025–Q1 2026 · verified

2026 speed-to-lead benchmark · verified

What is specific to this page.

Evidence
Kind Claim Check it against
WorkflowExtraction into defined fields is checkable because each value sits beside the source it came from, whereas a summary replaces the document with an account that cannot be verified at a glance or consumed by a downstream step.Comparing the time to verify an extracted field against verifying a summarised claim about the same document.
SoftwareDeterministic validations between extraction and use — total against line sum, date plausibility, supplier reference resolution, currency membership — each catch a class of extraction error mechanically at low cost.Applying each validation to a batch of extracted documents and counting the errors surfaced.
ConstraintRetaining and linking the original is what makes the extraction verifiable by a person or an auditor and recoverable when extraction proves to have been wrong for a period, so a system that discards the source has converted a document into an assertion.Attempting to verify an extracted value where the source document was not retained.
ResponseExtraction accuracy depends principally on document consistency rather than on general capability, so per-sender measurement identifies where manual handling should remain in a way an overall rate cannot.Computing extraction error rates grouped by sender across a month of documents.

Each row would be wrong on another industry's page. Where a sourced figure exists it is in the table above instead; these are the constraints that shape the work and do not happen to be numbers.

Start with the measurement.

Reading about a benchmark is not the same as knowing your own number. The audit produces yours, measured rather than estimated.

Get your growth audit

$497 · delivered in 5 business days · credited against month one