OCR processing on the local network

Scans. Sorted. Searchable.

Scanworker is local OCR software for duplex scans. It joins front and back scans into one document, corrects orientation, and publishes an early searchable PDF. A slower local VLM analysis proceeds independently in the background.

Scanworker project support

Why Scanworker exists

I use a Brother MFC-J6710DW. Its automatic document feeder scans only one side, so I scan every front page first, turn the paper stack over, and then scan every back page. Brother does not provide a feature for this particular model that automatically combines those two simplex passes, using this method, into a correctly ordered duplex document. I created Scanworker to fill that gap. You can find more about me and other projects at ralfhartmann.dev.

Built for a real needScanworker validates both scan passes, reverses the back-page order, interleaves front and back pages, and then adds text recognition. What began as a solution for my own scanner can therefore also work with other devices whose ADF scans only one side.

Duplex scanning with a simplex ADF

The paper stack is passed through the scanner twice. After the first pass, the complete stack is removed from the output tray without changing its order, rotated flat by 180 degrees in the plane of the supporting surface, and loaded again. The stack is not flipped over, and individual sheets are not turned. Both PDF files are then arranged by Scanworker into the correctly ordered duplex document.

  1. In a side view, a paper stack enters the ADF's upper input tray from the right.
    Front-page loadingThe aligned document stack is placed in the simplex ADF.
  2. In a side view, half of the first scanned stack protrudes from the inner output area; a horizontal arrow indicates removal.
    Stack removalAfter the first pass, the complete stack is removed from the output tray without changing its order.
  3. The upper paper stack is shown before rotation with an orange marker at the lower right. Below it, a 180-degree arrow points to the rotated stack, whose marker is now at the upper left. The suggested text lines are also rotated by 180 degrees.
    Flat 180° stack rotationThe complete stack is rotated flat on the supporting surface. It is neither flipped nor reordered, and individual sheets are not turned. This rotation gives the fast combined PDF the expected orientation; pages can additionally be rotated automatically during the later OCR pass.
  4. In a side view, the stack rotated flat by 180 degrees enters the upper ADF tray from the right again.
    Back-page loadingThe rotated stack is placed back into the ADF for the second scan pass.
  5. In a side view, half of the complete stack protrudes from the inner output area after the second pass.
    Second pass completeAfter the second pass, the stack is back in the output tray. Scanworker now combines both scan PDFs.

Choose the matching inputPut the two PDFs from a simplex two-pass scan in DUPLEX_INPUT_DIR; Scanworker pairs and interleaves them automatically. Put an already-complete PDF from a duplex ADF, or an ordinary single-sided scan, in SINGLE_INPUT_DIR. It enters text recognition automatically without being merged.

Two lanes, one document

The fast result never waits for compute-intensive analysis. Each processing stage publishes its own clearly identified artifact.

  1. Combination

    Two matching simplex PDFs are validated, the back pages are reversed, and both files are interleaved into the correct order. If automatic detection fails, exactly two input PDFs can be selected and re-enqueued manually—the job starts only when their page counts match.

  2. Fast OCR

    OCRmyPDF and Tesseract recognize German and English, correct upside-down documents, and run multiple independent page jobs concurrently.

  3. Background VLM

    PaddleOCR-VL analyzes each page's layout and text. Normalized text and layout data remain available as a separate JSON artifact.

Text-recognition technologies

The three local OCR paths have distinct roles and therefore publish different results.

OCRmyPDF and Tesseract

OCRmyPDF drives Tesseract for German and English. It deskews pages, corrects their orientation, and creates the positionally accurate text layer in the quickly available searchable PDF.

PaddleOCR-VL

The local VLM service analyzes each page and provides Markdown, normalized text, layout blocks, and bounding boxes as JSON. The “enhanced PDF” text layer currently continues to come from Tesseract.

Apple Vision

Optionally, Poppler renders PDF pages at 300 dpi; Tesseract detects suitable regions, while Apple Vision recognizes text and positions. pikepdf writes them as an invisible text layer into a separate Vision PDF.

Results as soon as they exist

The operations interface updates queue, page number, progress, OCR service, and estimated completion time live over WebSocket. More pairs can be enqueued while a job is running. Current and pending work is separated from history; pending and historical entries can be removed without deleting their PDFs.

Combined PDF

Available immediately after safe merging, before a text layer has been added.

Searchable PDF

Positionally accurate Tesseract text layer, automatic orientation correction, and content-derived filename.

VLM text and layout

Paddle Markdown, normalized document text, page provenance, bounding-box data, and audited cross-page word joins.

A deliberate technical boundaryVLM text is not yet emitted as a positionally accurate PDF text layer. Bounding boxes are retained for that future enhancement; the current “enhanced PDF” still uses Tesseract for its text layer.

Designed for quiet operation

Scanworker runs as unprivileged systemd services. State and queue data are durable, publication is atomic, and interrupted jobs can resume after restart.

Debian and macOS

Reproducible .deb and PKG/DMG builds, semantic versions, and separate development and release artifacts.

Local first

OCR and document analysis stay on the trusted network. Remote VLM targets require explicit permission.

PDF library at a glance

Page count and preview appear directly in the library. Clicking any column header sorts ascending or descending; a small arrow shows the direction, with newest PDFs first by default.

Voluntary support

Scanworker project support

Scanworker is free software and will remain free of charge. Its development can be supported voluntarily through PayPal. A payment does not create any additional rights, services, or benefits.