OCRmyPDF and Tesseract
OCRmyPDF drives Tesseract for German and English. It deskews pages, corrects their orientation, and creates the positionally accurate text layer in the quickly available searchable PDF.
OCR processing on the local network
Scanworker is local OCR software for duplex scans. It joins front and back scans into one document, corrects orientation, and publishes an early searchable PDF. A slower local VLM analysis proceeds independently in the background.
I use a Brother MFC-J6710DW. Its automatic document feeder scans only one side, so I scan every front page first, turn the paper stack over, and then scan every back page. Brother does not provide a feature for this particular model that automatically combines those two simplex passes, using this method, into a correctly ordered duplex document. I created Scanworker to fill that gap. You can find more about me and other projects at ralfhartmann.dev.
Built for a real needScanworker validates both scan passes, reverses the back-page order, interleaves front and back pages, and then adds text recognition. What began as a solution for my own scanner can therefore also work with other devices whose ADF scans only one side.
The paper stack is passed through the scanner twice. After the first pass, the complete stack is removed from the output tray without changing its order, rotated flat by 180 degrees in the plane of the supporting surface, and loaded again. The stack is not flipped over, and individual sheets are not turned. Both PDF files are then arranged by Scanworker into the correctly ordered duplex document.
Choose the matching inputPut the two PDFs from a simplex two-pass scan in DUPLEX_INPUT_DIR; Scanworker pairs and interleaves them automatically. Put an already-complete PDF from a duplex ADF, or an ordinary single-sided scan, in SINGLE_INPUT_DIR. It enters text recognition automatically without being merged.
The fast result never waits for compute-intensive analysis. Each processing stage publishes its own clearly identified artifact.
Two matching simplex PDFs are validated, the back pages are reversed, and both files are interleaved into the correct order. If automatic detection fails, exactly two input PDFs can be selected and re-enqueued manually—the job starts only when their page counts match.
OCRmyPDF and Tesseract recognize German and English, correct upside-down documents, and run multiple independent page jobs concurrently.
PaddleOCR-VL analyzes each page's layout and text. Normalized text and layout data remain available as a separate JSON artifact.
The three local OCR paths have distinct roles and therefore publish different results.
OCRmyPDF drives Tesseract for German and English. It deskews pages, corrects their orientation, and creates the positionally accurate text layer in the quickly available searchable PDF.
The local VLM service analyzes each page and provides Markdown, normalized text, layout blocks, and bounding boxes as JSON. The “enhanced PDF” text layer currently continues to come from Tesseract.
Optionally, Poppler renders PDF pages at 300 dpi; Tesseract detects suitable regions, while Apple Vision recognizes text and positions. pikepdf writes them as an invisible text layer into a separate Vision PDF.
The operations interface updates queue, page number, progress, OCR service, and estimated completion time live over WebSocket. More pairs can be enqueued while a job is running. Current and pending work is separated from history; pending and historical entries can be removed without deleting their PDFs.
Available immediately after safe merging, before a text layer has been added.
Positionally accurate Tesseract text layer, automatic orientation correction, and content-derived filename.
Paddle Markdown, normalized document text, page provenance, bounding-box data, and audited cross-page word joins.
A deliberate technical boundaryVLM text is not yet emitted as a positionally accurate PDF text layer. Bounding boxes are retained for that future enhancement; the current “enhanced PDF” still uses Tesseract for its text layer.
Scanworker runs as unprivileged systemd services. State and queue data are durable, publication is atomic, and interrupted jobs can resume after restart.
Reproducible .deb and PKG/DMG builds, semantic versions, and separate development and release artifacts.
OCR and document analysis stay on the trusted network. Remote VLM targets require explicit permission.
Page count and preview appear directly in the library. Clicking any column header sorts ascending or descending; a small arrow shows the direction, with newest PDFs first by default.
Voluntary support
Scanworker is free software and will remain free of charge. Its development can be supported voluntarily through PayPal. A payment does not create any additional rights, services, or benefits.