What OCR Gets Wrong on Real Freight Documents
Short answer
Character-level OCR reads text; extraction must also identify fields and document structure. Faxing, stamps, handwriting and phone photos can degrade results, while template-dependent tools may struggle with new layouts. Even correctly extracted documents can disagree. Test difficult scans, unfamiliar layouts, unreadable fields and cross-document comparison before choosing a freight-document tool.
Why can template-based extraction fail on freight documents?
Template-based extraction works by learning where a field sits on a page. That is viable when documents come from a handful of sources in a stable format.
Freight does not work that way. A mid-sized forwarder receives documents from hundreds of shippers, each with their own layout, and new shippers arrive constantly. A template per vendor is a configuration task that never finishes.
Where can character-level OCR break down?
| Condition | What goes wrong |
|---|---|
| Faxed or photocopied documents | Banding and dropout remove or distort characters |
| Stamps overlapping text | Characters under the stamp are misread or lost |
| Handwritten annotations | Treated as noise or misread as printed text |
| Phone photographs | Skew, shadow and uneven focus degrade recognition |
| Carbon copies | Low contrast produces unreliable character confidence |
| Multi-language documents | Characters outside the expected set are dropped |
What problem do extraction accuracy figures fail to capture?
A tool can report high field-level accuracy and still leave you exposed, because extraction accuracy measures whether a field was read correctly from one document. It says nothing about whether that document agrees with the others on the same shipment.
A commercial invoice reading 1,240 kg and a bill of lading reading 1,420 kg can both be extracted with complete accuracy. Both numbers are correctly read. The shipment still has a problem, and no extraction metric will surface it.
What should you test before buying a document tool?
-
Give it your worst scan, not a clean PDF. A stamped fax with handwriting in the margin is the real test.
-
Give it a document from a vendor it has never seen, with no configuration.
-
Give it two documents from one shipment where a field disagrees, and see whether anything notices.
-
Check what happens when a field is unreadable — does it say so, or does it guess?
-
Run twenty documents, not two. One good result is an anecdote.
The third test distinguishes extraction from cross-document validation, an important difference for teams preparing customs entries.
For the document workflow described here, see See it compare two documents.