What is OCR?
OCR (optical character recognition) reads the characters in an image and turns them into text a computer can work with. Text in scanned PDFs and photos is just an image, so it can’t be copied or searched. After OCR, you can edit and reuse it as text.
When OCR helps
- Extract text from scanned paper documents to reuse it
- Turn addresses and phone numbers on business cards or flyers into text
- Quote part of a book or document
- Make image-only documents editable in Word
- Extract vertical Japanese text
Tips for better accuracy
- Photograph the text straight on, in good light, so it’s sharp
- Avoid shadows, glare and camera shake
- If the text is small, take a closer photo of that part
- Choose “Japanese (vertical)” for vertical documents
- Handwriting and decorative typefaces are harder to read
Recognized text can contain mistakes. Always compare numbers, names and similar-looking characters (such as 力 and カ) with the original image.
Using the recognized text
You can correct the recognized text right on the page. Use “Copy all” to paste it into another app, or save it as a text file (.txt) or Word file (.docx). For PDFs, the text is separated by page.
If your PDF already has text
PDFs made from Word or Excel already contain text data, so you don’t need OCR. PDF to Word extracts their text faster and more accurately.
Your images are never uploaded
Most OCR services send your image to a server to read it. This tool loads the recognition program and data into your browser and processes everything on your device. Documents with personal information are never sent anywhere.
That means speed depends on your device. PDFs with many pages take a while, so you may want to take out just the pages you need with Split PDF first.