In healthcare, an OCR is judged on what it flags, not on what it reads

What an OCR must guarantee on a prescription or a lab request: line-by-line reading, a confidence score per value, referential matching, traceability. And why human review stays the rule.

Author and Co-Founder at Koncile
By 
Jules Ratier
Last updated: 
September 30, 2026
 - 
8
 min read

Every OCR vendor advertises high accuracy rates, and most of them deliver. On a health document, that is not the right buying criterion. What counts is what the tool does when it is not sure, what it lets you verify, and what it guarantees about the data it handles.

On an invoice, a misread line costs a correction. On a prescription, it costs a dose. The document is no longer a data source, it is a care instruction, and the software that reads it enters a chain where the error cannot always be caught downstream.

That is what makes the usual buying criteria insufficient. The question is not what an OCR can read, every vendor advertises high rates and most of them deliver. The question is what it does when it is not sure, what it lets you verify, and what it guarantees about the data it handles. Here is what we have learned deploying intelligent document processing pipelines for laboratories, hospitals, home healthcare providers and pharmacies.

What an OCR must guarantee on a health document

Read the line, not just the header

Plenty of tools can extract a patient name, a date and a prescriber. That is the easy part: those fields are few, always in the same place, and an error there is visible.

The work starts in the body of the document. A prescription is six, ten, sometimes twenty treatment lines, each with its molecule, its dose, its form, its posology and its duration. A lab request is a list of tests that has to come back in full, with none lost and none invented. A tool that returns the header and summarises the body is of no use to you.

The test is simple: count the lines on the document, count the ones that come out. On a batch of handwritten prescriptions processed at a French medical biology group, 301 tests came back out of 304 present. That is the ratio to look at, not the character recognition rate.

Hold the doses, the units and the forms

The same treatment can exist in ten different strengths. An extracted value therefore means nothing separated from its unit, and a confusion between two presentations of the same molecule produces a line that is perfectly formed and perfectly wrong.

Require the tool to return the unit with the value every time, to normalise when you ask it to, and to flag the cases where the unit is absent from the document rather than inferring it.

Handle handwriting

Handwriting is not a marginal case in this sector. Prescriptions written by hand, notes in the margin, photos taken on a phone by the patient, poor-quality scans: that is the daily reality of most processing chains.

A classic OCR, which recognises characters without understanding what it reads, breaks down on these documents. Semantic reading, where the model interprets the structure and the meaning of the document, holds up far better, because it can lean on context to choose between two possible readings. On the batch mentioned above, 23 of the 24 prescriptions were handwritten.

Handwriting recognition is a capability to test explicitly on your own documents, never to assume because it appears on a datasheet.

Printed prescription carrying a handwritten frequency note, correctly detected by the OCR
The handwritten note « tous les mercredis » (every Wednesday), added by hand on a printed prescription, comes back in the extraction. Without it, the other values could all be correct and the line would still be unusable: a treatment without its frequency is not a prescription.

Match every line against a referential

This is the requirement most often forgotten, and the one that sinks the most projects. No product code travels on a prescription. There is a label, written by hand by a prescriber with their own habits, and it has to be tied to an entry in your nomenclature.

Extracting without matching moves the manual work, it does not remove it. Your teams will go from typing to searching a referential, which is barely faster. It is a data matching problem as much as an extraction one, and it has methods of its own.

Data matching: acronyms from a biology request tied to the full labels and to the laboratory codes
Data matching in practice. « PTH » on the prescription, Parathormone # sérum and the laboratory code on the way out. Neither the label nor the code appears on the document.

Give a confidence score per value, not per document

This is the distinction that decides everything else. A document-level score tells you nothing you can act on: it tells you that something may be wrong somewhere.

A score per extracted value tells you which line to look at. That is the difference between rereading a whole document and checking three fields.

Confidence scores shown value by value on an extracted document, next to the document average
A score per value, not a score per document. That is the difference between rereading a whole prescription and checking three fields.

Trace who saw what

On these documents, traceability is not a comfort feature. You have to be able to say, months later, which document was processed, by which pipeline, with what result, and who accessed it.

Why a very good accuracy rate does not remove the need for human review

What the remaining percentage represents

Our deployments run at around 99 % of lines correctly extracted. That is a good level, and that is exactly why it deserves a close look.

A provider processing a million pages a year at 99 % produces ten thousand wrong lines over the year. In this sector, a wrong line is not a statistical anomaly: it is a posology, a patient, a billed procedure. The missing percentage never becomes negligible because it is small, it becomes dangerous because it is invisible.

No serious vendor should sell you the removal of review on this kind of document. We do not, and our customers in the sector do not ask us to: they fit extraction in as one more brick in a verification chain that remains theirs.

Not all errors cost the same

A misspelled prescriber name gets corrected as you go. A confused dose, a test tied to the wrong code, a sample prescribed in the wrong medium are not corrected the same way.

That is why the right indicator is not an average rate but the distribution of errors by criticality. Ask your vendor to produce it on your own documents.

What happens when these requirements are not set

A model that does not know would rather answer

This is the most dangerous behaviour, and the least known to buyers. Faced with information that is absent or illegible, a generative model tends to fill in rather than stay silent.

Two examples seen in production. When a prescriber omits the place of issue, the model goes and finds the town in the practice letterhead and returns it as if it appeared on the prescription. When the title reads « Madame », it infers the patient's gender, which is wrong in a share of cases and particularly problematic for transgender people.

The most telling case came from an identity document check. The system returned this: the documents appear consistent with one another, with no obvious sign of fraud, but the absence of legible information on the health insurance card makes it impossible to verify the holder's identity. A reassuring sentence about a file it had not been able to verify. A rushed operator remembers the first half.

A health OCR has to be configured to prefer silence over invention, and to surface explicitly what it could not read.

The referential that breaks silently

A case that took us several rounds to understand. At one customer, one test never came back. No error, no alert, simply a line missing every time.

The cause was not in the model. The code for that test in the customer's referential was NA, and the import pipeline read it as the value « not applicable ». The entry disappeared before matching even started. The workaround came down to one letter: the code became NAS, and that is the form it takes in the matching screenshot above.

Another customer was seeing unexplained coding gaps. The nomenclature file loaded for the demonstration was a simplified version, missing part of the entries.

In both cases the extraction engine was working perfectly. It was the chain around it that produced the defect, and no accuracy rate would have revealed it.

The gap between the document and the data entry

The benchmark people forget to take is the existing process. When we compared automatically extracted data against the manual entries of a home healthcare provider, roughly one file in forty showed a gap between the prescription and what had been recorded in the system.

That figure is not an argument against the teams, it is the reality of a repetitive task carried out under volume pressure. What it mainly says is that the question is not « is automation reliable enough », but « compared to what ».

Human review is not a failure of automation, it is its condition

A verification brick, not a replacement

The healthcare organisations that succeed do not plug in an OCR to remove a step. They insert it into a document workflow that already had several checks, so that human verification applies to already structured data rather than to a PDF.

The gain is not the disappearance of review. It is that review finally bears on something comparable, line by line, instead of forcing the operator to reconstruct the document in their head.

Target the checks instead of rereading everything

This is where the per-value confidence score takes on its operational meaning. You set a threshold, the lines below it go to an operator, the rest pass.

Setting that threshold is a business decision, not a technical one. It depends on what an error costs in your chain, and it has to be adjustable without a redeployment.

At a provider processing tens of thousands of prescriptions a month, a full review took 3 minutes 40 per document. With prior extraction and targeted checking, processing drops to 25 seconds followed by a verification. Over a year's volume, the saving exceeds 9,000 hours, close to six full-time equivalents. Those people have not disappeared, they have stopped retyping.

Integration best practices

Build a test set without exporting patient data

This is the first concrete obstacle, and it is legal before it is technical. In healthcare you often do not have the right to send real documents to a vendor for a trial, which makes the usual advice of « test on your real documents » inapplicable as it stands.

Three routes work. Build a corpus of fabricated documents reproducing your hard cases, which is quicker than it sounds. Anonymise upstream, provided the anonymisation covers every re-identifying element and not just the name. Or set up a patient consent procedure, which takes time but allows real documents.

Plan this step from the scoping phase. A project that discovers the problem at pilot stage loses several weeks.

Test on your full referential, not the demo one

That is the lesson from the case above. Load your entire nomenclature, with its exotic codes, its duplicates and its legacy entries. That is where the defects show up, never on a fifteen-line extract prepared for the presentation.

Measure the gap with the existing process before deciding

Before setting an accuracy target, measure the discrepancy rate of your current process. You will get an honest basis for comparison, and often a surprise.

Define the escalation thresholds

Decide explicitly which values never pass without human review, whatever the score. Doses and identities usually belong in that category.

Separate environments and version the models

Your extraction rules will evolve. They have to be able to do so without touching production, and you have to be able to roll back. A vendor offering neither a separate test environment nor versioning forces you to choose between standing still and taking the risk.

Scale up progressively

One batch, then a limited flow, then the switch. Each stage needs its own pass criterion, defined before you start.

What compliance adds

The questions your buyers will ask anyway

Security teams and data protection officers in the sector have professionalised their vendor reviews. Questionnaires now cover GDPR, security and, for the sector's financial players, DORA. Some groups run them through dedicated platforms and treat the compliance check as a non-negotiable prerequisite to any selection.

Anticipate them. The points that come back every time:

  • where the data is hosted, and whether the host is certified for health data (HDS in France)
  • what the retention period is, and whether it is configurable
  • whether there is a way to delete a specific document, and to verify that it is gone
  • whether processed data is used to train the models
  • whether the contract carries the processor clauses under Article 28 of the GDPR

To these is now added the question of the framework applicable to the AI itself. Depending on how it is used, an extraction system inserted into a care pathway may fall under additional obligations under the AI Act. The subject still appears rarely in vendor reviews, and it is better examined at scoping than discovered in committee.

The national insurance number, a case of its own

Counter-intuitive but important: in France, the social security number is not legally health data. It falls under a regime of its own, more constraining in some respects, because its structure is public knowledge and it identifies a person unambiguously. The permitted uses are listed in the texts, and what is not listed is not allowed.

If your pipeline extracts that number, check that your purpose falls within the cases provided for. It is a frequent blind spot in digitisation projects.

Verify rather than believe

A good practice observed at a demanding customer: keep the identifier returned for every document sent, precisely so you can go and check, later, that the data really was deleted. An advertised deletion capability is only worth what you can audit.

What it looks like in production

The orders of magnitude we see at our customers in the healthcare sector, on production pipelines: around 99 % of lines correctly extracted, on documents that are largely handwritten; processing time per document divided by several orders of magnitude, targeted review replacing full review; and a measured gap of about one file in forty between the documents and the previous manual entries.

These figures do not say that humans have become unnecessary. They say where humans now belong.

For more on the tools available, see our comparison of OCR tools for healthcare.

Frequently asked questions
Can an OCR read a handwritten prescription?

Yes, provided it relies on semantic reading rather than plain character recognition. The model uses context to choose between two possible readings, which lets it handle documents a human would struggle to decipher. A confidence score must accompany every value to flag the uncertain cases.

Do you need health data hosting certification to process health documents?

In France, HDS certification covers the hosting of health data. It is required of the host, and the controller must be able to prove that its hosting provider holds it. A software vendor generally acts as a processor under Article 28 of the GDPR, and that is the basis on which its contract should be examined. The two requirements add up, they do not replace one another.

How do you measure an OCR's real accuracy on your own documents?

By counting the expected lines and the correctly returned lines, on a representative batch that includes your hard cases, and by classifying the errors by criticality. An overall rate with no distribution supports no decision.

Can human review be removed?

No, and no vendor should promise it on health documents. What changes is the nature of the review: instead of rereading every document in full, teams check the values flagged as uncertain.

What happens when the model is not sure?

It depends on how it is configured, and that is a question to ask explicitly. A poorly configured model fills missing information by inference. A correctly configured one leaves the field empty and flags it.

How do you integrate extraction into an existing business application?

Through an OCR API. Documents leave your platform's backend, results come back as structured JSON with the fields, the confidence scores and the referential matches. The end user never sees the extraction tool and verification happens in your own interface.

Assessing an OCR on your own documents

If your teams are retyping prescriptions, lab requests or patient administrative documents, and you want to know what automated extraction would give at your organisation, we configure your fields, your rules and your nomenclature. You judge on your own corpus, and you can start by looking at our extraction for healthcare.

Book a demo

The agents that automate your documents
Get ahead on automation. See how Koncile can simplify your operations.
Discover Koncile
Discover Koncile
Our latest ARTICLES

Real life insights on document automation

All our ressources
All our ressources
Document processing and fraud detection: build or buy?
FEATURE

Document processing and fraud detection: build or buy?

Extracting and structuring document data, enriching it, detecting fraud: with AI, building your own end-to-end solution has never looked so simple. Buying off the shelf still often turns out to be the better option. Here are the thirty-nine difficulties to expect before you start.

Read the article
In healthcare, an OCR is judged on what it flags, not on what it reads
FEATURE

In healthcare, an OCR is judged on what it flags, not on what it reads

Every OCR vendor advertises high accuracy rates, and most of them deliver. On a health document, that is not the right buying criterion. What counts is what the tool does when it is not sure, what it lets you verify, and what it guarantees about the data it handles.

Read the article
Top Healthcare OCR Tools: HIPAA Compliance and What the Clinical Evidence Shows
FEATURE

Top Healthcare OCR Tools: HIPAA Compliance and What the Clinical Evidence Shows

Amazon Textract has been formally HIPAA-eligible since October 2019. Most healthcare OCR comparisons never mention that, or check which of the other four tools on their list actually offer a signed BAA. Here is a rebuild that checks compliance status directly, and looks at what the real clinical trials on ambient AI scribes actually found, not just what the marketing claims.

Read the article
Best Accounting Software for Sole Proprietors and Freelancers in 2026
FEATURE

Best Accounting Software for Sole Proprietors and Freelancers in 2026

A freelancer sending a dozen invoices a month can blow through Xero's cheapest plan by week three, and a $19-a-month FreshBooks account caps out at five clients. The number on a pricing page is rarely the number you actually pay. Here is a real comparison of QuickBooks, Xero, Wave, and FreshBooks built for a business of one, not a growing team.

Read the article
Hebrew OCR: what works on real business documents
Analysis

Hebrew OCR: what works on real business documents

Most OCR vendors say they support Hebrew. Far fewer support Hebrew handwriting, and none of them tells you that 91% character accuracy can mean one word in two is wrong.

Read the article