HomePDF ToolsPDF OCR

🔍 PDF OCR

Extract text from scanned PDFs and images using Optical Character Recognition (OCR). Supports multiple languages. No upload — runs 100% in your browser.

ℹ️ How it works: Tesseract OCR (Google's open-source OCR engine) runs in your browser to read text from scanned pages. On first use it downloads the language data file (~10 MB) — this is normal and only happens once. Works on scanned PDFs and image files (JPG, PNG, TIFF).
🔍
Drop PDF or image here
PDF, JPG, PNG, TIFF — up to 50 MB
📄
🇺🇸 English
🇮🇳 Hindi
🇫🇷 French
🇩🇪 German
🇪🇸 Spanish
🇵🇹 Portuguese
🇨🇳 Chinese (Simplified)
🇯🇵 Japanese
🔒 100% Free · No Login · File Never Leaves Your Browser · Powered by Tesseract.js
📖 How to Use
1
Upload File
Drop a scanned PDF or image file (JPG, PNG, TIFF)
2
Choose Language
Select the language of the text in your document
3
Copy or Download
Copy the extracted text or download as a .txt file
❓ Frequently Asked Questions
Why does it download data on first use?+
Tesseract needs language training data to recognize text. The English data is about 10 MB and is cached in your browser — subsequent uses are instant without downloading again.
Does it work on normal PDFs with text?+
For PDFs with real selectable text, use PDF to Word instead — it's faster and more accurate. OCR is for scanned or image-only PDFs where the text is a picture.
Are my files sent to a server?+
No. Tesseract.js runs fully in your browser using WebAssembly. Your file never leaves your device.