❓ Frequently Asked Questions
Why does it download data on first use?+
Tesseract needs language training data to recognize text. The English data is about 10 MB and is cached in your browser — subsequent uses are instant without downloading again.
Does it work on normal PDFs with text?+
For PDFs with real selectable text, use PDF to Word instead — it's faster and more accurate. OCR is for scanned or image-only PDFs where the text is a picture.
Are my files sent to a server?+
No. Tesseract.js runs fully in your browser using WebAssembly. Your file never leaves your device.