Skip to content

Glossary · Capture

OCR

Technology that converts a scanned document or an image into machine-readable text so it can be searched, extracted and archived.

Also known as: Optical Character Recognition · text recognition

What is OCR?

OCR (Optical Character Recognition) recognises the characters in an image — a scan, a photo of a contract, a non-text PDF — and produces a text layer. Without OCR, an “image” PDF is just a picture: you can see it, but you can’t search it, extract from it or index it. In business, OCR is the gateway to any serious document automation.

Why OCR matters in business

Incoming documents (e-mail, post, supplier portals) often arrive as scans. Without OCR, teams re-key dates, amounts and references — a source of errors and delays.

Searching archives

Finding a contract by SIRET company number, clause or amount requires indexed text, not just a file name.

Preparing for extraction

OCR feeds classification and data extraction: end dates, parties, amounts.

Less re-keying

Less manual typing = fewer discrepancies between the PDF and your business records.

Common pitfalls

  • Confusing OCR with understanding

    OCR reads characters; it does not decide the document type or its legal validity. Validation remains a human task.

  • Neglecting image quality

    Blurry photos, folded pages or very low-resolution scans reduce the recognition rate — and with it the whole downstream pipeline.

OCR in DocPilot

DocPilot applies OCR to incoming documents (upload, e-mail, photo) to make their content searchable and prepare assisted extraction before validation.

  • Multi-channel capture followed by automatic OCR
  • Indexed text for document search
  • Foundation for extraction and approval workflows

Take action

Make your scans usable

Turn on DocPilot capture: OCR, a structured record and an approval flow as soon as documents arrive.