Skip to content
ENRU
All articles

How to test AI before trusting it with documents

Any model gets things wrong sometimes. What matters is how often, and whether you catch it. Here’s how to run a small pilot that shows whether to go further.

A small robot checking a stack of documents with a magnifying glass

Language models can now do things that until recently needed a person. They pull details out of scanned documents, answer questions from internal policies and sort incoming requests by topic. They also get things wrong, calmly and confidently, without a hint of doubt. So there’s no point asking whether a model will make mistakes. It will. What you need to know is how often, in which places, and whether you’ll spot the error before it ends up in your books.

Pick one task

Pilots that start with “let’s bring in AI” nearly always disappoint, because nobody has said what success would look like. What works is a narrow, slightly boring goal. “Extract the details from incoming invoices.” “Answer questions about returns using only the returns policy.” Something like “an assistant for the team” is too vague to measure.

With one task you get one result you can measure. You can check it, compare it with how the work is done today, and make a decision based on numbers instead of impressions.

Give it an exam

Before you hand documents to a model, test it the way you’d test a new hire. For that you need a reference set of 50–100 real examples where the right answer is already known. That might be invoices an accountant has already checked, or customer questions paired with the answers your most experienced person would give.

Without a test like this, any opinion on quality is just a feeling. With it, you’re looking at numbers, and there’s much less to argue about.

What to measure

The first thing people look at is accuracy, meaning how many answers match the reference. A single overall figure can hide a lot, though. A model might read amounts perfectly and still mix up dates, so count accuracy separately for each field.

The second thing, and probably the most important, is whether the model can admit it’s unsure. A good system says “not sure” when it isn’t. A doubtful answer can be sent to a person to check. A wrong answer given with full confidence slips straight through. We’d take a system that flags its doubts over one with slightly better accuracy that never flags anything.

The third is time. How long does it take someone to check the model’s output, compared with doing the job by hand? If checking takes as long as typing, you’ve gained nothing, however good the other numbers look.

Keep a person in the loop

A well-built system changes what an employee does instead of replacing them. Typing data in becomes checking it. For that to work, the interface has to show where each value came from, highlight anything doubtful and keep a record of every correction.

That correction log gets less attention than it deserves. It’s your quality control, and over time it’s also the material that helps the model make fewer mistakes.

Deciding what counts as success

Write down your success criteria before the pilot starts, so you can’t adjust them afterwards. For example: at least 95 % accuracy on the key fields, at least nine out of ten doubtful cases correctly flagged, and document processing time cut in half. The numbers will be different for every task. What matters is agreeing on them up front.

If the pilot misses its targets, you’ve still learned something useful. You found out on a small budget, before rolling anything out across the company.

Where your data goes

If the documents contain personal data or commercial secrets, the first question is where they’re being sent. With a cloud model, you hand the data to a provider under a contract. A model on your own server keeps the data inside the company, but you’ll need the hardware to run it. Both options work, and your security requirements should decide between them.

Treat it like a new hire

Test AI the way you’d test a new employee. Give it a clear task, set an exam using real examples, let it say “I don’t know”, and keep an experienced person next to it. A pilot on a single task, with a reference set and targets agreed in advance, tells you whether to keep going before you’ve spent serious money.

All articles

Does your business need its own app?
Bot or mini app: taking bookings and orders in Telegram
Connecting your website, CRM and accounting so nobody copies data by hand
Taking a project over from your previous developer without losing anything

Working on something similar? Let’s talk

Where to reply
What you’re interested in · optional
Project estimate

Send us your project

Describe the job in your own words and attach whatever you have: a brief, mockups, spreadsheets or a link to your site. We’ll take a look and tell you roughly what it could cost and where it makes sense to start.

Files · optional

Drop files here or PDF, Word, Excel, PowerPoint, TXT, PNG, JPG, WEBP. Up to 5 files, 20 MB in total
Where to reply

@ittectorittector@gmail.com