OCR can extract text from a scanned PDF with 99% accuracy — or produce garbled nonsense. The difference comes down to a handful of factors you can control before you even scan.
OCR (Optical Character Recognition) accuracy is the percentage of characters correctly identified from a scanned image. A 99% accuracy rate on a 500-word page (~3,000 characters) still means around 30 errors — enough to require careful proofreading.
Modern OCR engines on clean, high-resolution scans typically achieve 98–99.5% accuracy. But accuracy drops sharply when conditions are poor.
| Condition | Typical Accuracy | Rating |
|---|---|---|
| Clean scan, 300+ DPI, white background | 98–99.5% | Excellent |
| Clean scan, 150–299 DPI | 90–97% | Good |
| Slightly skewed page (<5°) | 85–95% | Acceptable |
| Low DPI (72–150), or blurry scan | 60–85% | Poor |
| Coloured or patterned background | 50–80% | Poor |
| Handwriting (printed block letters) | 60–80% | Poor |
| Cursive handwriting | 20–50% | Very poor |
300 DPI is the minimum for reliable OCR. At 72 DPI (screen resolution), individual characters blur together and error rates spike. Always scan at 300 DPI or higher for documents you plan to OCR.
A page tilted even 3–5° causes OCR to misread characters. Most modern OCR tools auto-deskew, but a seriously skewed scan (from hand-holding a phone) significantly reduces accuracy.
Black text on white paper gives the best contrast. Coloured paper, watermarks, or graphical backgrounds make it hard for OCR to distinguish letters from background noise.
Standard fonts (Arial, Times New Roman, Helvetica) are recognised at near-perfect accuracy. Decorative, handwritten, or very small fonts (<8pt) drop accuracy considerably.
Creased, torn, stained, or faded documents introduce noise that OCR interprets as characters. Flatten pages before scanning for best results.
Our PDF to Word tool uses OCR to extract text from scanned PDFs.
Try PDF to Word →OCR accuracy is the percentage of characters correctly recognised from a scanned image. Modern OCR on clean 300 DPI scans typically achieves 98–99.5% accuracy.
Standard OCR is poor at handwriting — typically 60–80% accuracy for printed block letters, and 20–50% for cursive. For best results, documents should be typed.
300 DPI is the minimum recommended for reliable OCR. 400–600 DPI gives the best results. Below 200 DPI, character recognition errors increase significantly.
Common causes: low scan resolution, skewed pages, blurry images, coloured backgrounds, unusual fonts, or damaged source documents.