OCR Scanner - Extract Text from Images and PDFs
Recognize Vietnamese and English text from photos, scans and PDF pages entirely in your browser, then export as TXT, Word or a searchable PDF.
Related tools
How it works
- 1
Add images or a PDF
Drop, paste or browse any photo format or a PDF. Each PDF page becomes its own page in the workspace.
- 2
Choose a language and recognize
Pick Vietnamese, English, or both, then recognize all pages. Processing runs locally using Tesseract.js.
- 3
Review side by side
Compare the recognized text against the source image for each page, and fix any mistakes directly in the text box.
- 4
Export your result
Download the combined text as TXT, a Word document, or a searchable PDF with an invisible, copyable text layer over the original image.
FAQ
- Are my images or PDFs uploaded anywhere?
- No. Recognition runs entirely in your browser using Tesseract.js. Only the OCR engine and language data (not your files) are downloaded from a CDN the first time you use it, then cached. A recovery copy of your current session is stored on this device and expires after 48 hours.
- Which languages are supported?
- Vietnamese, English, or both combined in a single pass, for documents that mix the two scripts.
- What does the searchable PDF export do?
- It draws the recognized words as invisible, selectable text directly over the original page image, so the PDF looks unchanged but its text can be selected, copied and searched.
- What are the limits?
- Images up to 30 MB, PDFs up to 50 MB, and up to 30 pages processed per PDF. Very large or low-quality scans take longer and may recognize less accurately.
- Can I fix mistakes in the recognized text?
- Yes. Each page's text is editable — correct it directly before exporting.
