How to test AI before trusting it with documents
Any model gets things wrong sometimes. What matters is how often, and whether you catch it. Here’s how to run a small pilot that shows whether to go further.

Language models can now do things that until recently needed a person. They pull details out of scanned documents, answer questions from internal policies and sort incoming requests by topic. They also get things wrong, calmly and confidently, without a hint of doubt. So there’s no point asking whether a model will make mistakes. It will. What you need to know is how often, in which places, and whether you’ll spot the error before it ends up in your books.
Pick one task
Pilots that start with “let’s bring in AI” nearly always disappoint, because nobody has said what success would look like. What works is a narrow, slightly boring goal. “Extract the details from incoming invoices.” “Answer questions about returns using only the returns policy.” Something like “an assistant for the team” is too vague to measure.
With one task you get one result you can measure. You can check it, compare it with how the work is done today, and make a decision based on numbers instead of impressions.
Give it an exam
Before you hand documents to a model, test it the way you’d test a new hire. For that you need a reference set of 50–100 real examples where the right answer is already known. That might be invoices an accountant has already checked, or customer questions paired with the answers your most experienced person would give.
Without a test like this, any opinion on quality is just a feeling. With it, you’re looking at numbers, and there’s much less to argue about.
What to measure
The first thing people look at is accuracy, meaning how many answers match the reference. A single overall figure can hide a lot, though. A model might read amounts perfectly and still mix up dates, so count accuracy separately for each field.
The second thing, and probably the most important, is whether the model can admit it’s unsure. A good system says “not sure” when it isn’t. A doubtful answer can be sent to a person to check. A wrong answer given with full confidence slips straight through. We’d take a system that flags its doubts over one with slightly better accuracy that never flags anything.
The third is time. How long does it take someone to check the model’s output, compared with doing the job by hand? If checking takes as long as typing, you’ve gained nothing, however good the other numbers look.
Keep a person in the loop
A well-built system changes what an employee does instead of replacing them. Typing data in becomes checking it. For that to work, the interface has to show where each value came from, highlight anything doubtful and keep a record of every correction.
That correction log gets less attention than it deserves. It’s your quality control, and over time it’s also the material that helps the model make fewer mistakes.
Deciding what counts as success
Write down your success criteria before the pilot starts, so you can’t adjust them afterwards. For example: at least 95 % accuracy on the key fields, at least nine out of ten doubtful cases correctly flagged, and document processing time cut in half. The numbers will be different for every task. What matters is agreeing on them up front.
If the pilot misses its targets, you’ve still learned something useful. You found out on a small budget, before rolling anything out across the company.
Where your data goes
If the documents contain personal data or commercial secrets, the first question is where they’re being sent. With a cloud model, you hand the data to a provider under a contract. A model on your own server keeps the data inside the company, but you’ll need the hardware to run it. Both options work, and your security requirements should decide between them.
Treat it like a new hire
Test AI the way you’d test a new employee. Give it a clear task, set an exam using real examples, let it say “I don’t know”, and keep an experienced person next to it. A pilot on a single task, with a reference set and targets agreed in advance, tells you whether to keep going before you’ve spent serious money.



