scanbutler
ghcr.io/tom-joad/scanbutler:latest
https://ghcr.io
bridge
sh
false
https://github.com/Tom-Joad/scanbutler
https://github.com/Tom-Joad/scanbutler
https://raw.githubusercontent.com/Tom-Joad/unraid-templates/main/templates/scanbutler.xml
https://raw.githubusercontent.com/Tom-Joad/unraid-templates/main/icons/scanbutler.png
Turns scanned paper into searchable PDFs, one per document, named by topic and date. The
Stacks input splits large scans (hundreds of pages, no separator sheets) into documents by
content. The Scanner input takes files from a document scanner, one document per file, and
gives them the same OCR and naming. Uses Mistral OCR (batch API) and a Mistral chat model;
the text layer comes from a fresh Tesseract pass. No web UI: progress is in the container
log, and a review.md per stack in the Work folder lists uncertain splits. An optional third
input only adds the text layer and uploads to Paperless-ngx. Optionally reports the queue to
a webhook such as Home Assistant. Needs a Mistral API key.
Productivity: Tools:
--security-opt no-new-privileges --memory=4g