
In healthcare, an OCR is judged on what it flags, not on what it reads
What an OCR must guarantee on a prescription or a lab request: line-by-line reading, a confidence score per value, referential matching, traceability. And why human review stays the rule.

Every OCR vendor advertises high accuracy rates, and most of them deliver. On a health document, that is not the right buying criterion. What counts is what the tool does when it is not sure, what it lets you verify, and what it guarantees about the data it handles.
On an invoice, a misread line costs a correction. On a prescription, it costs a dose. The document is no longer a data source, it is a care instruction, and the software that reads it enters a chain where the error cannot always be caught downstream.
That is what makes the usual buying criteria insufficient. The question is not what an OCR can read, every vendor advertises high rates and most of them deliver. The question is what it does when it is not sure, what it lets you verify, and what it guarantees about the data it handles. Here is what we have learned deploying intelligent document processing pipelines for laboratories, hospitals, home healthcare providers and pharmacies.
What an OCR must guarantee on a health document
Read the line, not just the header
Plenty of tools can extract a patient name, a date and a prescriber. That is the easy part: those fields are few, always in the same place, and an error there is visible.
The work starts in the body of the document. A prescription is six, ten, sometimes twenty treatment lines, each with its molecule, its dose, its form, its posology and its duration. A lab request is a list of tests that has to come back in full, with none lost and none invented. A tool that returns the header and summarises the body is of no use to you.
The test is simple: count the lines on the document, count the ones that come out. On a batch of handwritten prescriptions processed at a French medical biology group, 301 tests came back out of 304 present. That is the ratio to look at, not the character recognition rate.
Hold the doses, the units and the forms
The same treatment can exist in ten different strengths. An extracted value therefore means nothing separated from its unit, and a confusion between two presentations of the same molecule produces a line that is perfectly formed and perfectly wrong.
Require the tool to return the unit with the value every time, to normalise when you ask it to, and to flag the cases where the unit is absent from the document rather than inferring it.
Handle handwriting
Handwriting is not a marginal case in this sector. Prescriptions written by hand, notes in the margin, photos taken on a phone by the patient, poor-quality scans: that is the daily reality of most processing chains.
A classic OCR, which recognises characters without understanding what it reads, breaks down on these documents. Semantic reading, where the model interprets the structure and the meaning of the document, holds up far better, because it can lean on context to choose between two possible readings. On the batch mentioned above, 23 of the 24 prescriptions were handwritten.
Handwriting recognition is a capability to test explicitly on your own documents, never to assume because it appears on a datasheet.

Match every line against a referential
This is the requirement most often forgotten, and the one that sinks the most projects. No product code travels on a prescription. There is a label, written by hand by a prescriber with their own habits, and it has to be tied to an entry in your nomenclature.
Extracting without matching moves the manual work, it does not remove it. Your teams will go from typing to searching a referential, which is barely faster. It is a data matching problem as much as an extraction one, and it has methods of its own.
.webp)
Give a confidence score per value, not per document
This is the distinction that decides everything else. A document-level score tells you nothing you can act on: it tells you that something may be wrong somewhere.
A score per extracted value tells you which line to look at. That is the difference between rereading a whole document and checking three fields.

Trace who saw what
On these documents, traceability is not a comfort feature. You have to be able to say, months later, which document was processed, by which pipeline, with what result, and who accessed it.
Why a very good accuracy rate does not remove the need for human review
What the remaining percentage represents
Our deployments run at around 99 % of lines correctly extracted. That is a good level, and that is exactly why it deserves a close look.
A provider processing a million pages a year at 99 % produces ten thousand wrong lines over the year. In this sector, a wrong line is not a statistical anomaly: it is a posology, a patient, a billed procedure. The missing percentage never becomes negligible because it is small, it becomes dangerous because it is invisible.
No serious vendor should sell you the removal of review on this kind of document. We do not, and our customers in the sector do not ask us to: they fit extraction in as one more brick in a verification chain that remains theirs.
Not all errors cost the same
A misspelled prescriber name gets corrected as you go. A confused dose, a test tied to the wrong code, a sample prescribed in the wrong medium are not corrected the same way.
That is why the right indicator is not an average rate but the distribution of errors by criticality. Ask your vendor to produce it on your own documents.
What happens when these requirements are not set
A model that does not know would rather answer
This is the most dangerous behaviour, and the least known to buyers. Faced with information that is absent or illegible, a generative model tends to fill in rather than stay silent.
Two examples seen in production. When a prescriber omits the place of issue, the model goes and finds the town in the practice letterhead and returns it as if it appeared on the prescription. When the title reads « Madame », it infers the patient's gender, which is wrong in a share of cases and particularly problematic for transgender people.
The most telling case came from an identity document check. The system returned this: the documents appear consistent with one another, with no obvious sign of fraud, but the absence of legible information on the health insurance card makes it impossible to verify the holder's identity. A reassuring sentence about a file it had not been able to verify. A rushed operator remembers the first half.
A health OCR has to be configured to prefer silence over invention, and to surface explicitly what it could not read.
The referential that breaks silently
A case that took us several rounds to understand. At one customer, one test never came back. No error, no alert, simply a line missing every time.
The cause was not in the model. The code for that test in the customer's referential was NA, and the import pipeline read it as the value « not applicable ». The entry disappeared before matching even started. The workaround came down to one letter: the code became NAS, and that is the form it takes in the matching screenshot above.
Another customer was seeing unexplained coding gaps. The nomenclature file loaded for the demonstration was a simplified version, missing part of the entries.
In both cases the extraction engine was working perfectly. It was the chain around it that produced the defect, and no accuracy rate would have revealed it.
The gap between the document and the data entry
The benchmark people forget to take is the existing process. When we compared automatically extracted data against the manual entries of a home healthcare provider, roughly one file in forty showed a gap between the prescription and what had been recorded in the system.
That figure is not an argument against the teams, it is the reality of a repetitive task carried out under volume pressure. What it mainly says is that the question is not « is automation reliable enough », but « compared to what ».
Human review is not a failure of automation, it is its condition
A verification brick, not a replacement
The healthcare organisations that succeed do not plug in an OCR to remove a step. They insert it into a document workflow that already had several checks, so that human verification applies to already structured data rather than to a PDF.
The gain is not the disappearance of review. It is that review finally bears on something comparable, line by line, instead of forcing the operator to reconstruct the document in their head.
Target the checks instead of rereading everything
This is where the per-value confidence score takes on its operational meaning. You set a threshold, the lines below it go to an operator, the rest pass.
Setting that threshold is a business decision, not a technical one. It depends on what an error costs in your chain, and it has to be adjustable without a redeployment.
At a provider processing tens of thousands of prescriptions a month, a full review took 3 minutes 40 per document. With prior extraction and targeted checking, processing drops to 25 seconds followed by a verification. Over a year's volume, the saving exceeds 9,000 hours, close to six full-time equivalents. Those people have not disappeared, they have stopped retyping.
Integration best practices
Build a test set without exporting patient data
This is the first concrete obstacle, and it is legal before it is technical. In healthcare you often do not have the right to send real documents to a vendor for a trial, which makes the usual advice of « test on your real documents » inapplicable as it stands.
Three routes work. Build a corpus of fabricated documents reproducing your hard cases, which is quicker than it sounds. Anonymise upstream, provided the anonymisation covers every re-identifying element and not just the name. Or set up a patient consent procedure, which takes time but allows real documents.
Plan this step from the scoping phase. A project that discovers the problem at pilot stage loses several weeks.
Test on your full referential, not the demo one
That is the lesson from the case above. Load your entire nomenclature, with its exotic codes, its duplicates and its legacy entries. That is where the defects show up, never on a fifteen-line extract prepared for the presentation.
Measure the gap with the existing process before deciding
Before setting an accuracy target, measure the discrepancy rate of your current process. You will get an honest basis for comparison, and often a surprise.
Define the escalation thresholds
Decide explicitly which values never pass without human review, whatever the score. Doses and identities usually belong in that category.
Separate environments and version the models
Your extraction rules will evolve. They have to be able to do so without touching production, and you have to be able to roll back. A vendor offering neither a separate test environment nor versioning forces you to choose between standing still and taking the risk.
Scale up progressively
One batch, then a limited flow, then the switch. Each stage needs its own pass criterion, defined before you start.
What compliance adds
The questions your buyers will ask anyway
Security teams and data protection officers in the sector have professionalised their vendor reviews. Questionnaires now cover GDPR, security and, for the sector's financial players, DORA. Some groups run them through dedicated platforms and treat the compliance check as a non-negotiable prerequisite to any selection.
Anticipate them. The points that come back every time:
- where the data is hosted, and whether the host is certified for health data (HDS in France)
- what the retention period is, and whether it is configurable
- whether there is a way to delete a specific document, and to verify that it is gone
- whether processed data is used to train the models
- whether the contract carries the processor clauses under Article 28 of the GDPR
To these is now added the question of the framework applicable to the AI itself. Depending on how it is used, an extraction system inserted into a care pathway may fall under additional obligations under the AI Act. The subject still appears rarely in vendor reviews, and it is better examined at scoping than discovered in committee.
The national insurance number, a case of its own
Counter-intuitive but important: in France, the social security number is not legally health data. It falls under a regime of its own, more constraining in some respects, because its structure is public knowledge and it identifies a person unambiguously. The permitted uses are listed in the texts, and what is not listed is not allowed.
If your pipeline extracts that number, check that your purpose falls within the cases provided for. It is a frequent blind spot in digitisation projects.
Verify rather than believe
A good practice observed at a demanding customer: keep the identifier returned for every document sent, precisely so you can go and check, later, that the data really was deleted. An advertised deletion capability is only worth what you can audit.
What it looks like in production
The orders of magnitude we see at our customers in the healthcare sector, on production pipelines: around 99 % of lines correctly extracted, on documents that are largely handwritten; processing time per document divided by several orders of magnitude, targeted review replacing full review; and a measured gap of about one file in forty between the documents and the previous manual entries.
These figures do not say that humans have become unnecessary. They say where humans now belong.
For more on the tools available, see our comparison of OCR tools for healthcare.
Assessing an OCR on your own documents
If your teams are retyping prescriptions, lab requests or patient administrative documents, and you want to know what automated extraction would give at your organisation, we configure your fields, your rules and your nomenclature. You judge on your own corpus, and you can start by looking at our extraction for healthcare.








