Dev.to WebDev ๐Ÿ›  Dev ๐Ÿ‘ 0 ๐Ÿ“– 1 min read

Free OCR sites cap you at 15 pages/hour, so I built an unlimited one that runs Tesseract in the browser

Every free OCR service I tried throttles you: 15 pages/hour, 50 pages/month, or a free tier that exists to push you to a paid plan. The cap makes sense for them, because server compute costs money on every page. The only

Every free OCR service I tried throttles you: 15 pages/hour, 50 pages/month, or a free tier that exists to push you to a paid plan. The cap makes sense for them, because server compute costs money on every page. The only genuinely unlimited free option has been raw Tesseract on the command line, which most people can't use.

So I ran the same engine in the browser instead:

  • Tesseract.js (Tesseract compiled to WebAssembly) runs on your own CPU, so there's no per-page cost and nothing to ration
  • Accepts JPG, PNG, WEBP, HEIC (converted on-device) and multi-page scanned PDFs, processed page by page
  • Text shows up per page as it finishes, not after the whole batch
  • Each page returns Tesseract's confidence score, and low-confidence pages get flagged instead of silently handing you garbled text
  • 6 languages for now: English, Spanish, French, German, Hindi, Portuguese
  • The language model downloads once (a few MB) and is cached. Your document itself is never uploaded, which you can verify in the Network tab

Limits worth stating up front: it's a print-text engine, so no handwriting, no table or form reconstruction, and blurry low-resolution scans will struggle.

๐Ÿ”— https://www.forgeplug.com/tools/ocr-text-extractor

Which languages would be most useful to add next?

๐Ÿ“ฐ Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes โ€” full credit and traffic to the original publisher.