byADIT Free software for engineers and students
EngiFree
OCRmyPDF — The Complete Guide: Making a Scanned Archive of Reports and Standards Searchable

OCRmyPDF — The Complete Guide: Making a Scanned Archive of Reports and Standards Searchable

OCRmyPDF is a command-line tool that runs optical character recognition (OCR) on a scanned PDF and places the recognised text beneath the image. The page looks exactly the same, but now you can search and copy from it — ideal for an archive of old geotechnical reports, standards and project files.

What you get

What to know

It is built on the Tesseract engine and supports more than 100 languages. There is no graphical interface, and on Windows you must install Tesseract and Ghostscript first. Licensed under MPL-2.0 — free for any use.

Related tools

The engine itself — Tesseract OCR. For scanning straight to searchable PDF — NAPS2. For a graphical OCR front end — gImageReader.

The bottom line

One command on a folder, and a whole scanned archive becomes searchable. Well worth the quarter of an hour it takes to install.

Further reading

מדריכי AI למהנדסים ולמשרד, ב-5 שפות: adit-ai.com ←