zgba Network

Your Extraction Is 98% Accurate. One Document in Three Still Needs a Human.

If you have ever built a document extraction demo, you know how good it feels. Ten sample invoices go into a model, twenty tidy JSON fields come out, and everyone in the room starts doing math on how many hours of data entry just disappeared. Then it goes live, and the hours mostly do not disappear. Not because the model got worse, but because the demo measured the wrong thing. I have ended up thinking about document pipelines less as a model problem and more as a validation and UI problem, and this post is the version of that argument I wish I had read before the first one. It is three jobs, not one “Read the document” is really three separate jobs. Classify. Work out what actually arrived: an invoice, a credit note, a statement that looks like an invoice, a photo of a delivery note, or three of those stapled into one PDF. A perfect extraction sent down the wrong path is still a bug. Extract. Pull out supplier, date, invoice number, line items, tax, total. This is the part every demo shows. Post. Write the result into the system that runs the business, matched to the right supplier or case, without creating duplicates. This is an integration project in disguise, and in my experience it eats more of the build than extraction does. There is also a question worth asking before any of that: does this document need to exist? A lot of paper survives only because nobody redesigned intake. If the information comes from people you can reach, capturing it as structured data at the source beats reading it back out of a PDF later. That is the idea behind Fortell AI, where we build voice and SMS systems that help Community Action Agencies handle intake in over 100 languages, so information arrives as data instead of paperwork somebody has to retype. Extraction is the right tool for documents you do not control: supplier invoices, third-party forms, anything sent in someone else’s format. 98% accurate is a per-field number Accuracy figures are almost always quoted per field. Your users experience the system per document. If each of twenty fields is right 98 times out of 100, the chance that every field on a document is right is 0.98 to the power of 20, which is about 0.67. So roughly one document in three has at least one wrong field, and nobody knows which third. In practice that means someone still checks everything, and the savings quietly evaporate. The number that actually decides whether the project was worth it is straight-through rate: the share of documents that land in the destination system correctly with nobody touching them. The second number is time to handle the rest: how long a flagged document takes to clear compared with typing it from scratch. Both need a baseline, which means timing the current manual process on a real week of real documents, ugly ones included. Without that, you cannot say afterwards whether you saved anything. Expect the rate to vary wildly by sender. A regular supplier with a clean generated PDF might go straight through nearly every time. A one-off scanned form with a stamp over the total might never go straight through. That is fine, as long as the system knows which is which. Do not ask the model how sure it is The obvious design is: ask the model for a confidence score, route anything low to a human. It is also the weakest one. Language models are not reliably calibrated about their own mistakes, and the errors that hurt are confident ones: a total read from the wrong column, day and month swapped, two digits transposed in an invoice number. The strong signal is checking extracted values against things you already know. Roughly, the shape I reach for looks like this (illustrative, not lifted from a specific codebase): type Flag = { field: string; reason: string }; function validateInvoice(inv: ExtractedInvoice, ctx: KnownData): Flag[] { const flags: Flag[] = []; // 1. The document’s own arithmetic const lineSum = inv.lines.reduce((s, l) => s + l.amount, 0); if (Math.abs(lineSum - inv.subtotal) > 0.01) flags.push({ field: “subtotal”, reason: “line items do not sum” }); if (Math.abs(inv.subtotal + inv.tax - inv.total) > 0.01) flags.push({ field: “total”, reason: “subtotal + tax != total” }); // 2. Master data you already trust const supplier = ctx.suppliers.get(inv.supplierTaxId); if (!supplier) flags.push({ field: “supplier”, reason: “unknown supplier” }); else if (inv.total > supplier.typicalMax * 3) flags.push({ field: “total”, reason: “far outside normal range” }); // 3. Duplicates if (ctx.seenInvoice(inv.supplierTaxId, inv.invoiceNumber)) flags.push({ field: “invoiceNumber”, reason: “already received” }); // 4. Format and range if (inv.date > ctx.today) flags.push({ field: “date”, reason: “in the future” }); return flags; } None of that is clever, and that is the point. A document that fails its own math gets flagged no matter how confident the model was. Every check is deterministic, testable, and owned by you rather than by a prompt. Two more rules I hold to. Keep provenance on every field. Extraction is not summarization. On LectureNotes AI, one of my own products, the whole point is turning a lecture into paraphrased takeaways and clean outlines. Extraction has zero room for paraphrase: the number in your database has to be the number on the page. Storing where on the page each value came from (page, bounding box) makes review faster and debugging possible. Some fields never post on their own. A document asking to change a supplier’s bank details is the textbook invoice fraud pattern. Text inside a document should never authorize anything, whatever your checks say. I wrote more about that line in Give an AI Agent Write Access One Verb at a Time, and it applies to every PDF a pipeline reads. The review screen is the actual product For the first months, a real share of documents will need a person. How fast that person clears them decides the savings more than any model choice. If checking a flagged document is slower than typing it, the project has failed while the model succeeded. What I think a review screen needs: Document and fields side by side, with the source of each value highlighted on the page. Checking becomes looking, not searching. This is where the provenance pays off. Only the doubtful fields flagged. If the arithmetic passed and the supplier matched, do not ask the reviewer to re-read everything. Asking people to verify everything trains them to verify nothing. Keyboard first. One document after another, a fix in a few keystrokes, no modal dialogs. Every correction stored. A fixed field is a free labelled test case. Collect them and you have a regression set to rerun whenever the model, the prompt, or a supplier’s layout changes. Same discipline I use for voice agents in There Is No Repro for a Phone Call. The queue also needs a named owner. A pile of flagged documents nobody is responsible for becomes a backlog, and a backlog of unposted invoices is worse than the manual process it replaced, because at least the manual process was visible. Rollout follows the same ladder as anything that writes to real records: run alongside the manual process and compare, then draft-and-approve, then let clean document types from clean senders post on their own, one type at a time. What actually makes it hard Five things move the effort more than the model does: Layout variety. Five suppliers with consistent invoices is a small project. Five hundred senders in five hundred layouts is a different one, because the long tail never stops producing new cases. Input quality. Generated PDFs are the best case. Scans, phone photos, faxes, stamps and handwriting each lower straight-through rate and raise review load. Languages. Each language needs its own test set and someone who can read it. Where it lands. Posting through a documented accounting API is simpler than posting into a custom system, and posting money is stricter than posting a case note. If it feeds a ledger, it inherits everything in A Stored Balance Is a Rumor. Sensitivity. ID documents, bank details and health data change who may see the review queue and how long originals are kept. Settle that before the build, not during the security review. On build versus buy: off-the-shelf tools are genuinely good for common documents into common systems, like receipts into mainstream accounting software. A custom pipeline earns its keep when the documents are specific to an industry, when validation depends on your own data, or when the destination is a system you built yourself. The short version The model reading the page is the least uncertain part of a document pipeline. What decides whether it works is deterministic validation against data you already trust, a review screen fast enough that people use it properly, and a hard line around the fields that always wait for a human. Measure straight-through rate, not accuracy, and time the manual baseline before you write a line of code. The longer client-facing version of this, with the questions I would ask any vendor before commissioning one, is on the Null Studio blog. If you have shipped one of these, I am curious what your straight-through rate looked like in month one versus month six, and which check caught the most real errors.

View original article