All Tools Merge PDF Split PDF Compress PDF PDF to Word Edit PDF Blog Corporate Solutions Contact Us
Back to Blog
September 2026  ·  5 min read  ·  PDF Guide

OCR Accuracy Guide: What Affects Scanned PDF Text Extraction

OCR can extract text from a scanned PDF with 99% accuracy — or produce garbled nonsense. The difference comes down to a handful of factors you can control before you even scan.

What Is OCR Accuracy?

OCR (Optical Character Recognition) accuracy is the percentage of characters correctly identified from a scanned image. A 99% accuracy rate on a 500-word page (~3,000 characters) still means around 30 errors — enough to require careful proofreading.

Modern OCR engines on clean, high-resolution scans typically achieve 98–99.5% accuracy. But accuracy drops sharply when conditions are poor.

OCR Accuracy by Scan Quality

ConditionTypical AccuracyRating
Clean scan, 300+ DPI, white background98–99.5%Excellent
Clean scan, 150–299 DPI90–97%Good
Slightly skewed page (<5°)85–95%Acceptable
Low DPI (72–150), or blurry scan60–85%Poor
Coloured or patterned background50–80%Poor
Handwriting (printed block letters)60–80%Poor
Cursive handwriting20–50%Very poor

The Biggest Factors That Affect OCR Quality

1. Scan resolution (DPI) — most important factor

300 DPI is the minimum for reliable OCR. At 72 DPI (screen resolution), individual characters blur together and error rates spike. Always scan at 300 DPI or higher for documents you plan to OCR.

2. Page skew and rotation

A page tilted even 3–5° causes OCR to misread characters. Most modern OCR tools auto-deskew, but a seriously skewed scan (from hand-holding a phone) significantly reduces accuracy.

3. Background colour and patterns

Black text on white paper gives the best contrast. Coloured paper, watermarks, or graphical backgrounds make it hard for OCR to distinguish letters from background noise.

4. Font type

Standard fonts (Arial, Times New Roman, Helvetica) are recognised at near-perfect accuracy. Decorative, handwritten, or very small fonts (<8pt) drop accuracy considerably.

5. Physical document condition

Creased, torn, stained, or faded documents introduce noise that OCR interprets as characters. Flatten pages before scanning for best results.

Convert Scanned PDF to Word Free

Our PDF to Word tool uses OCR to extract text from scanned PDFs.

Try PDF to Word →

Tips for Best OCR Results

Tip: Always scan at 300 DPI minimum. If your scanner offers 400 or 600 DPI, use it for documents with small text or fine detail.
Tip: Place documents flat on the scanner bed rather than photographing them with a phone. Phone photos introduce lens distortion and uneven lighting that hurts OCR accuracy.
Tip: After OCR conversion, always proofread the output — especially numbers, special characters (%, $, @), and proper nouns which are most commonly misread.
Tip: If OCR results are poor, try scanning the original document again at higher DPI rather than re-running OCR on an already-converted file.

Frequently Asked Questions

What is OCR accuracy?

OCR accuracy is the percentage of characters correctly recognised from a scanned image. Modern OCR on clean 300 DPI scans typically achieves 98–99.5% accuracy.

Can OCR read handwriting?

Standard OCR is poor at handwriting — typically 60–80% accuracy for printed block letters, and 20–50% for cursive. For best results, documents should be typed.

What DPI is needed for good OCR?

300 DPI is the minimum recommended for reliable OCR. 400–600 DPI gives the best results. Below 200 DPI, character recognition errors increase significantly.

Why does OCR produce garbled text?

Common causes: low scan resolution, skewed pages, blurry images, coloured backgrounds, unusual fonts, or damaged source documents.

Related Tools

Processing your PDF...

Your file is ready