← All notes

Compare OCR Engines: Tesseract, Surya and pdfplumber in One Docker Image

Take one invoice and test it on four OCRs.

The first reads only the text already typed into the PDF, in under a second. The second squints at the pixels and tells you how sure it is about every word. The third understands the page like a human: where the table starts, what comes first, what's a header. The fourth lives in the cloud and takes its time.

You get four different texts. Which one is right?

I kept asking myself that question, project after project. So I put all four OCRs in one box.

ocrHub is a free, open-source Docker image that runs several OCR engines (Tesseract, pdfplumber, Surya, Datalab) on the same document, so you can compare their results side by side. Its purpose is comparison, and speed comes second.

ocrHub dashboard for OCR engine comparison: the original page next to the pdfplumber, Tesseract, Surya and Datalab results, with boxes drawn on each

One test page, four engines, side by side. pdfplumber took 0.3 s, Tesseract 3.8 s and Datalab 10.9 s. pdfplumber and Surya recover the sales table as a grid, while Tesseract returns it as plain lines of text.

Why is there no single best OCR engine?

Because "OCR" covers very different jobs. A born-digital PDF, a photographed receipt and a dense newspaper page each favor a different engine. The only reliable way to choose is to test on your own files. That's what ocrHub is for.

How do Tesseract, Surya, pdfplumber and Datalab differ?

Engine Best for What you get Typical speed (CPU, my tests)
pdfplumber Born-digital PDFs Text layer, fonts, tables Under 1 second
Tesseract Clean scans and images Text, per-word confidence, line boxes 2-20 s per page
Surya Complex layouts Layout regions, reading order, tables About 3 min for a 2-page PDF; needs about 6 GB RAM
Datalab (cloud) Layout and tables without local load Layout blocks, HTML tables 10 s to 2.5 min; paid, with a free monthly tier

These timings come from my own tests on an older MacBook Pro running Docker Desktop with an 8 GB memory limit, so treat them as rough. Newer hardware will be faster.

Which OCR engine should I pick?

  • Born-digital PDF (text you can select): pdfplumber. It reads the text layer directly, so it is fast and exact.
  • Clean scan or photo: Tesseract. It is light, and its per-word confidence shows where to double-check.
  • Complex layout, tables or reading order: Surya locally, or Datalab if you accept a cloud service.

Still unsure? Run your own file through all of them and look at the differences.

How do I run an OCR comparison with Docker?

One command, no building:

docker run -p 8000:8000 -v "$PWD/data:/home/ocrhub/data" ghcr.io/misky8/ocrhub:latest

Open localhost:8000, upload an image or PDF, tick the engines, and view the results side by side or as an overlay. This image includes Tesseract and pdfplumber (about 640 MB). To add Surya and Datalab, build from source with ENGINES=surya,datalab. Full instructions are in the GitHub README.

Can an AI agent use OCR through MCP or an API?

Yes. The same engines are available through an HTTP API (POST /ocr) and through an MCP server, so an agent like Claude Code can call OCR directly:

claude mcp add --transport http ocrhub http://localhost:8000/mcp/

Is it free and self-hosted?

Yes. ocrHub is MIT-licensed and runs on CPU in one container, and your files stay on your machine. The only cloud engine is Datalab, and only if you turn it on.

Try it, star it, or tell me what broke: github.com/MiSky8/ocrHUB

Need OCR wired into a document pipeline for your company? At Nexi8 invoices, contracts and scans become structured data for a knowledge base, reports and internal AI agents.

Which document would you test first: a scanned receipt, or the one that gives you nightmares?