Searching archives
Finding a contract by SIRET company number, clause or amount requires indexed text, not just a file name.
Glossary · Capture
Technology that converts a scanned document or an image into machine-readable text so it can be searched, extracted and archived.
Also known as: Optical Character Recognition · text recognition
OCR (Optical Character Recognition) recognises the characters in an image — a scan, a photo of a contract, a non-text PDF — and produces a text layer. Without OCR, an “image” PDF is just a picture: you can see it, but you can’t search it, extract from it or index it. In business, OCR is the gateway to any serious document automation.
Incoming documents (e-mail, post, supplier portals) often arrive as scans. Without OCR, teams re-key dates, amounts and references — a source of errors and delays.
Finding a contract by SIRET company number, clause or amount requires indexed text, not just a file name.
OCR feeds classification and data extraction: end dates, parties, amounts.
Less manual typing = fewer discrepancies between the PDF and your business records.
OCR reads characters; it does not decide the document type or its legal validity. Validation remains a human task.
Blurry photos, folded pages or very low-resolution scans reduce the recognition rate — and with it the whole downstream pipeline.
DocPilot applies OCR to incoming documents (upload, e-mail, photo) to make their content searchable and prepare assisted extraction before validation.
To learn more about the product and the guides.
Step-by-step methods to make real progress.
DocPilot blog posts on the same topic.
Related pages to refine your research.
Take action
Turn on DocPilot capture: OCR, a structured record and an approval flow as soon as documents arrive.