All apps · 0 apps
scanbutler
Docker app from Tomjoad's Repository
Overview
Readme
View on GitHubScanbutler
Turn scanned paper into searchable PDFs, one per document, named after their content. It runs as a Docker container that watches up to three folders.
Stacks. Drop a scan of a whole pile of paper, with hundreds of pages and no separator sheets. The container finds where each document begins from the content alone and splits the scan:
stacks/inbox/Household/stack-01.pdf (500 pages)
↓
stacks/output/Household/Blood count 2026-09-30.pdf
stacks/output/Household/CT report chest 2026-09-30.pdf
stacks/output/Household/Electricity bill 2026-08-14.pdf
...
Scanner. Point a document scanner that saves to a network share at the second inbox. Each file there is one document. It gets the same OCR and naming, but is never split:
scanner/inbox/20261002_141503.pdf
↓
scanner/output/Insurance renewal notice 2026-09-28.pdf
This suits Fujitsu (now Ricoh) ScanSnap scanners such as the iX1600 well. When they scan straight to a network folder (SMB), they save image-only PDFs: the OCR is done by the ScanSnap software on a computer. Point the scanner at this inbox, and the files come out searchable and named without any PC running.
Paperless (optional). Files dropped here only get the Tesseract text layer and are then uploaded to Paperless-ngx. Paperless, or an AI tagger working with it such as Zettelrobbe, takes care of the title, tags and correspondent:
paperless/inbox/scan.pdf → text layer → Paperless-ngx document #1234
Titles are written in the language of each document unless you set
TITLE_LANGUAGE.
Under the hood:
- ocrmypdf with Tesseract adds a fresh, invisible text layer, so every output PDF is searchable.
- Mistral OCR reads stacks through Mistral's batch API, at half the regular price. Its structured text makes splitting more reliable. Scanner files are read from the Tesseract layer instead, which costs nothing and adds no waiting time.
- A Mistral chat model decides where documents begin and names them.
Contents
- System requirements
- Quick start
- How it works
- Reviewing and correcting splits
- Configuration
- Languages
- Queue webhook (Home Assistant)
- Spending limit and paused processing
- Choosing the text source
- Paperless-ngx input
- Rate limits and cost
- Privacy
- Unraid
- Development
- Contributing
System requirements
| Minimum | Recommended | |
|---|---|---|
| Memory for the container | 1 GB | 4 GB |
| CPU | 64-bit x86 (amd64) or ARM (arm64) | 4 or more cores |
| Disk, work folder | a few GB free | several GB free for 500-page stacks |
| Software | Docker 20.10 or newer | |
| Network | HTTPS to api.mistral.ai; optionally Paperless and Home Assistant on the LAN |
The container adapts to its memory. With OCRMYPDF_JOBS=auto, the default,
it runs as many OCR pages in parallel as its memory limit allows: 1 GB as a
base plus 0.75 GB per page. It never runs more than one page per CPU the
container may use, and Docker's --cpus and --cpuset-cpus limits are
respected. Less memory therefore makes processing slower, never makes it
fail.
All inputs share that budget, and scanner and Paperless files go first:
- A file gets one job per page, as many as are free. A two-page scan takes two jobs and leaves the rest to others.
- When jobs are scarce, waiting scanner and Paperless files are served before stacks. The same goes for requests to Mistral.
- Stacks get their text layer in pieces of 20 pages. A scan therefore waits for one piece at most, even while a 500-page stack is running.
- The scanner and Paperless inputs work on several files at once.
Beyond a certain point, more memory no longer helps: the CPUs become the
limit. Full speed needs about 1 GB + 0.75 GB × CPU cores, so 5.5 GB for 6
cores, or 7 GB for 8. The startup log shows the outcome:
"memory_gb":6.0,"cpus":6,"ocr_jobs_total":6.
Measured with a worst-case file: four A4 pages, each holding a colour image
at 1550 dpi. The images are downsampled to 600 dpi first, see
OCRMYPDF_MAX_IMAGE_DPI.
| Container memory | Parallel pages | Peak memory | Time |
|---|---|---|---|
| 768 MB | 1 | 625 MB | 53 s |
| 4 GB | 4 | 1.64 GB | 24 s |
For scale: 156 ordinary scanned pages took about 4.5 minutes on a 12-thread desktop CPU. Expect a low-power NAS CPU to be several times slower. That is fine for a folder watcher, but a 500-page stack may then take hours.
Synology and other NAS systems: in Container Manager, set the memory limit
under Resources to at least 1 GB, or 4 GB if available. NAS models with
32-bit ARM CPUs (armv7) are not supported; check your model's CPU
architecture. Map /data to a shared folder and /config to a folder for
the container's own data, and set PUID/PGID to a user that may write
there, as with any linuxserver.io container.
Quick start
You need Docker and a Mistral API key.
git clone https://github.com/Tom-Joad/scanbutler.git
cd scanbutler
cp .env.example .env # set MISTRAL_API_KEY, and DATA_PATH, PUID, PGID, TZ if needed
docker compose up -d --build
docker compose logs -f
The prebuilt image ghcr.io/tom-joad/scanbutler can replace the
local build, see docker-compose.yml.
On first start, the container creates this layout under DATA_PATH:
stacks/ inbox/ output/ archive/ failed/
scanner/ inbox/ output/ archive/ failed/
paperless/ inbox/ archive/ failed/ (only with PAPERLESS_URL set)
paperless-2/ inbox/ archive/ failed/ (only with PAPERLESS_2_TOKEN set)
Its own data (plans, caches, review files) goes to CONFIG_PATH, mounted
at /config.
Put PDFs into stacks/inbox/ or scanner/inbox/, optionally in
sub-folders. Sub-folders are mirrored in output/. A file is picked up once
it has not changed for STABLE_SECONDS (60 s by default), so copying a large
scan over the network is safe.
To process a single file once, outside the inboxes:
docker compose exec scanbutler scanbutler process /data/some.pdf --profile stacks --folder "Household"
Commands run with docker exec act as the abc user, so their files get
the same owner as everything else.
How it works
OCR (stacks by default). Pages go to Mistral OCR in chunks of 50 through the batch API. A job takes minutes instead of seconds and costs half as much. Job ids and results are stored as soon as they exist, so an interrupted run never pays for a page twice. Uploads and results are deleted from Mistral's file storage afterwards. A chunk that fails inside a batch is retried directly.
OCR_MODE=directskips the batch API.Scanner files skip this step by default. Their text comes from the Tesseract layer that step 5 adds anyway. Either input can be switched with
STACKS_TEXT_SOURCE/SCANNER_TEXT_SOURCE; see Choosing the text source.Blank pages. A page is dropped when its image shows almost no ink, typically the back of a duplex scan. The check uses the image itself, not the OCR text, because OCR models sometimes invent whole paragraphs on an empty page.
Boundaries (stacks only). A chat model reads overlapping windows of 12 pages and decides, page by page, whether a new document starts. It looks at letterheads, salutations, headings, dates, layout changes and sentences that run across pages. Each page's verdict comes from the window where it had the most context on both sides. Explicit "page k of n" markers override the model. A second copy of a document becomes a file of its own, for example
... (2).pdf.Naming. Each document gets a short title of 1 to 6 words, naming its type and the detail that sets it apart, plus the date it is about. For reports and lab results, that is the examination or sampling date. For letters, it is the issue date.
Text layer. ocrmypdf adds Tesseract's text, choosing the mode per file:
- Pure scans (no text at all):
--force-ocr. Pages are deskewed, the image Tesseract reads is cleaned, and poor scans are upsampled to 300 dpi. - Tagged PDFs, which carry a logical structure tree: kept as they are. These are born-digital files such as office documents or online bank statements. Their text is the original, and re-OCR would discard the structure.
- Other PDFs that already contain text, such as scans that carry the
scanner's own OCR:
--redo-ocr. An old OCR layer is replaced, while real digital text stays as it is. Forcing OCR would turn such pages into pictures of themselves, many times the size. There is no image cleaning in this mode: ocrmypdf rasterizes at the resolution of the sharpest image on a page, and cleaning a page with a high-resolution logo at that size can exhaust the memory.
Images sharper than 600 dpi, such as a high-resolution logo, are first downsampled with Ghostscript. Text and vector graphics stay untouched. Without this, ocrmypdf would rasterize the whole page at the image's resolution. Every mode also runs within limits, so one odd page can't exhaust the server. Tesseract sees at most
OCRMYPDF_MAX_OCR_MPIXELSper page, and each page and each run has a time limit. The visible page is never downsampled.If a mode fails, a simpler one is tried before the file counts as failed. The last resort OCRs only pages without text and does no image processing. The log and
.error.txtshow ocrmypdf's actual error message. The image ships Tesseract'stessdata_bestmodels for German and English. They are slower than the defaults, but hold up much better on poor scans. Other languages are downloaded on first use, see Languages.- Pure scans (no text at all):
Output. Each document's pages are cut from the searchable scan into
<input>/output/<sub-folder>/<title> <date>.pdf. The PDF's title and subject metadata are set as well.
The original file then moves to archive/. A file that fails moves to
failed/ together with an .error.txt. Move it back into inbox/ to retry:
OCR that was already paid for is reused.
To check that a PDF really has a text layer, open it and search for a word
with Ctrl+F, or select the text. In the document's font list, the invisible
layer shows up as GlyphLessFont.
Reviewing and correcting splits
Without separator sheets, splitting cannot be perfect. Every input file gets
a work directory, work/<input>/<sub-folder>/<file name>-<hash>/. It holds:
review.md: every document with its page range. A document is flagged ⚠ when it starts mid-document (for example on "page 3 of 5"), when it is a single, nearly empty page, or when the model was unsure.decisions.json: the model's verdict and reason for every page.plan.json: the split plan, meant to be edited by hand. Page numbers are 1-based ranges such as"4-6, 9".
To fix a split, edit plan.json: change pages, or merge and split entries.
Set title to "" to have the title and date generated again. Then run:
docker compose exec scanbutler scanbutler rebuild "stacks/Household/stack-01-1a2b3c4d"
The files listed in written_files are replaced. Nothing is OCR'd again.
Configuration
All settings are environment variables. .env.example lists
them with comments.
Container (linuxserver.io conventions)
The image is built on linuxserver.io's
Debian base and behaves like their containers: it starts as root, gives the
abc user the IDs below and runs Scanbutler as abc.
| Parameter | Default | Purpose |
|---|---|---|
-e PUID |
911 |
User ID that owns new files; use the owner of your data folder (Unraid: 99) |
-e PGID |
911 |
Group ID for new files (Unraid: 100) |
-e UMASK |
022 |
Permission mask for new files; 002 makes them writable for the group |
-e TZ |
Etc/UTC |
Time zone, e.g. Europe/Berlin |
-v /config |
Work folder: plans, caches, review files, the Paperless upload ledger, temporary page images and downloaded languages | |
-v /data |
The inputs: stacks/, scanner/, paperless/, paperless-2/ |
Don't run the container with --user or --init: the base image's init
system (s6-overlay) has to start as root and as process 1. Docker mods
(DOCKER_MODS) work as with any linuxserver.io image.
Mistral
| Variable | Default | Purpose |
|---|---|---|
MISTRAL_API_KEY |
— | Required |
MISTRAL_LLM_MODEL |
mistral-large-latest |
Model for splitting and naming |
MISTRAL_OCR_MODEL |
mistral-ocr-latest |
OCR model |
MISTRAL_MAX_RPS |
1 |
Requests per second across all workers; set it below your account's limit |
OCR_MODE |
batch |
batch (half price, minutes) or direct (full price, seconds) |
BATCH_POLL_SECONDS |
15 |
How often a running batch job is checked |
BATCH_MAX_WAIT_HOURS |
24 |
A job still running after this is cancelled; its chunks are retried directly |
PAUSE_RETRY_MINUTES |
30 |
Probe interval while paused, see Spending limit |
MISTRAL_API_BASE |
https://api.mistral.ai/v1 |
API endpoint |
MISTRAL_TIMEOUT |
300 |
Seconds per request |
Folders and inputs
| Variable | Default | Purpose |
|---|---|---|
DATA_DIR |
/data |
Parent of the default folders below |
STACKS_DIR / SCANNER_DIR |
$DATA_DIR/stacks, $DATA_DIR/scanner |
Root of each input; inbox/, output/, archive/ and failed/ live below it |
STACKS_ENABLED / SCANNER_ENABLED |
true |
Switch an input off |
STACKS_TEXT_SOURCE / SCANNER_TEXT_SOURCE |
mistral / tesseract |
Text for splitting and naming: mistral (Mistral OCR) or tesseract (free, from the text layer). The scanner falls back to mistral when OCRMYPDF_ENABLED=false |
WORK_DIR |
/config |
OCR results, plans and review files |
WORK_RETENTION_DAYS |
30 |
Days after which the work folder of a processed file is deleted; rebuild works until then. A rebuild restarts the count. 0 keeps everything |
TMPDIR |
$WORK_DIR/tmp |
Temporary page images, several GB for a large stack; on disk, not in RAM |
POLL_INTERVAL |
30 |
Seconds between inbox checks |
STABLE_SECONDS |
60 |
A file must stay unchanged this long before it is picked up. A PDF that isn't completely written (no %%EOF at its end), or an empty one, waits up to 10 minutes longer and then moves to failed/. Any file still waiting after 10 minutes is named once in the log as file still waiting, with the reason |
Paperless-ngx input
| Variable | Default | Purpose |
|---|---|---|
PAPERLESS_URL |
— | Base URL of Paperless-ngx, e.g. http://paperless:8000; setting it enables the input |
PAPERLESS_TOKEN |
— | API token of the Paperless user that should own the documents |
PAPERLESS_DIR |
$DATA_DIR/paperless |
Root of the input; inbox/, archive/ and failed/ live below it |
PAPERLESS_TAGS |
— | Comma-separated tag ids to add on upload, e.g. 3,7 |
PAPERLESS_TEXT_SOURCE |
tesseract |
mistral replaces the document's content in Paperless with Mistral OCR's text, tables included; see below |
PAPERLESS_MAX_WAIT_MINUTES |
30 |
How long to wait for Paperless to consume a file before trying again later |
PAPERLESS_2_TOKEN |
— | API token of a second Paperless user; setting it enables paperless-2/inbox/, whose uploads belong to that user; see A second Paperless user |
PAPERLESS_2_URL |
PAPERLESS_URL |
Paperless URL for the second input |
PAPERLESS_2_DIR |
$DATA_DIR/paperless-2 |
Root of the second input |
PAPERLESS_2_TAGS |
— | Comma-separated tag ids for uploads through the second input |
PAPERLESS_2_TEXT_SOURCE |
tesseract |
As PAPERLESS_TEXT_SOURCE, for the second input |
PAPERLESS_SHARE_TAGS |
false |
true removes the owner from every tag, so all users see it; needs a superuser token; see Shared tags, correspondents and document types |
PAPERLESS_SHARE_CORRESPONDENTS |
false |
The same for correspondents |
PAPERLESS_SHARE_DOCUMENT_TYPES |
false |
The same for document types |
PAPERLESS_SHARE_TAGS_MINUTES |
1 |
How often they are checked, for all three kinds |
PAPERLESS_SHARE_TAGS_READONLY |
— | Comma-separated tag names that keep their owner and are only visible to other users, e.g. ai-processed |
Naming
| Variable | Default | Purpose |
|---|---|---|
TITLE_LANGUAGE |
each document's language | For example English or German |
FILENAME_PATTERN |
{title} {date} |
{title} is required; {date} is YYYY-MM-DD |
NO_DATE_LABEL |
undated |
Replaces {date} when no date is found |
Splitting and blank pages
| Variable | Default | Purpose |
|---|---|---|
BOUNDARY_WINDOW / BOUNDARY_STEP |
12 / 6 |
Pages per model request, and how far each window moves on |
REVIEW_CONFIDENCE |
0.75 |
Below this model confidence, a document is flagged |
DROP_BLANK_PAGES |
true |
false keeps blank pages with the document before them |
BLANK_MAX_INK_PERCENT |
0.2 |
Pages with less visible ink than this are blank |
BLANK_MAX_CHARS |
15 |
Pages with no image and at most this much text are blank |
METADATA_MAX_CHARS |
24000 |
Text per document sent for naming; longer documents are shortened in the middle |
Text layer
| Variable | Default | Purpose |
|---|---|---|
OCRMYPDF_ENABLED |
true |
false keeps the scan's own text layer, if any |
OCRMYPDF_LANGUAGES |
deu+eng |
Tesseract languages, joined with +, e.g. deu+eng+fra; see Languages |
TESSDATA_URL |
tessdata_best 4.1.0 on GitHub |
Where languages that aren't built in are downloaded from; a mirror must serve the same files |
OCRMYPDF_JOBS |
auto |
Parallel OCR pages, shared by all inputs; auto follows the memory and CPU limits, see System requirements |
OCRMYPDF_MAX_IMAGE_DPI |
600 |
Images sharper than this are downsampled before OCR; plenty for text, and it bounds each page's memory. 0 disables it |
OCRMYPDF_EXTRA_ARGS |
— | Appended to the ocrmypdf call |
OCRMYPDF_MAX_OCR_MPIXELS |
50 |
Larger page images are downsampled for OCR only. Peak memory is about jobs × 16 bytes × this value; A4 at 600 dpi is about 35 |
OCRMYPDF_PAGE_TIMEOUT |
300 |
Seconds Tesseract may spend on one page; the page is kept either way |
OCRMYPDF_FILE_TIMEOUT_MINUTES |
120 |
An ocrmypdf run taking longer is stopped, and the next mode is tried |
OCRMYPDF_SKIP_BIG_MPIXELS |
200 |
Last-resort mode only: pages above this are kept without new OCR |
Throughput, webhook and logging
| Variable | Default | Purpose |
|---|---|---|
OCR_CHUNK_PAGES |
50 |
Pages per OCR request |
OCR_CONCURRENCY / LLM_CONCURRENCY |
3 / 4 |
Parallel requests; MISTRAL_MAX_RPS still applies |
QUEUE_WEBHOOK_URL |
— | Report queue counts here, see below |
QUEUE_WEBHOOK_CHECK_SECONDS |
10 |
How often the queue is counted |
QUEUE_WEBHOOK_HEARTBEAT_SECONDS |
300 |
Resend interval without changes |
LOG_LEVEL |
INFO |
Logs are JSON lines on stdout. Every line logged while a file is being worked on names it in source (folder and name below the inbox) |
Languages
deu, eng and osd (page orientation) are built into the image. Any
other language in OCRMYPDF_LANGUAGES is downloaded at startup, once:
- Models come from
tessdata_bestat the pinned tag4.1.0. Every download is checked against a SHA-256 list shipped in the image; a file that doesn't match is discarded. - They are kept in
/config/tessdataand survive image updates. Delete a file there to download it again. - Codes are Tesseract's three-letter ones:
fra,ita,spa,pol,tur,chi_simand so on, 124 in total. A misspelt code stops the container with a hint, e.g.unknown language 'ger' (did you mean 'deu'?). - Without internet access, a language that isn't downloaded yet is left out
with a warning, and the rest keeps working; the built-in languages always
do. A local mirror can stand in via
TESSDATA_URL.
The start of the log says what is in use:
{"event":"languages ready","languages":"deu+eng+fra","downloaded":["fra"],"cached":[],"missing":[]}
More languages make OCR slower; list only the ones your documents use.
Queue webhook (Home Assistant)
With QUEUE_WEBHOOK_URL set, the container POSTs the queue state as JSON.
It sends whenever a number changes, and again every
QUEUE_WEBHOOK_HEARTBEAT_SECONDS, so the receiver catches up after a
restart. The payload holds counts only, never file names:
{
"queued": 3, "waiting": 2, "processing": 1, "failed": 0,
"paused": false, "pause_reason": null, "paused_since": null,
"profiles": {
"stacks": {"waiting": 1, "processing": 1, "failed": 0},
"scanner": {"waiting": 1, "processing": 0, "failed": 0}
}
}
queuediswaiting + processing.processingcan be more than 1: the scanner and Paperless inputs work on several files at once.failedcounts the PDFs in thefailed/folders.profileshas one entry per enabled input;paperlessappears only whenPAPERLESS_URLis set,paperless-2only whenPAPERLESS_2_TOKENis.paused,pause_reasonandpaused_sinceare always present. The last two arenullunless processing is paused.pause_reasonis Mistral's raw error text, up to 300 characters.paused_sinceis an ISO 8601 timestamp in UTC.
Every send and every failure appears in the container log, with the error text but never the URL. An unchanged error is repeated at most once per heartbeat interval. If the receiver is unreachable, processing carries on.
For Home Assistant, a trigger-based template entity reads the webhook. Pick a
long random webhook_id, because anyone who knows it can post to it:
template:
- triggers:
- trigger: webhook
webhook_id: scan-splitter-queue-CHANGE-ME
allowed_methods: [POST]
local_only: true
sensor:
- name: Scan-Splitter queue
unique_id: scan_splitter_queue
state: "{{ trigger.json.queued }}"
unit_of_measurement: files
attributes:
waiting: "{{ trigger.json.waiting }}"
processing: "{{ trigger.json.processing }}"
stacks_waiting: "{{ trigger.json.profiles.stacks.waiting | default(0) }}"
scanner_waiting: "{{ trigger.json.profiles.scanner.waiting | default(0) }}"
- name: Scan-Splitter failed
unique_id: scan_splitter_failed
state: "{{ trigger.json.failed }}"
unit_of_measurement: files
binary_sensor:
- name: Scan-Splitter paused
unique_id: scan_splitter_paused
state: "{{ trigger.json.paused | default(false) }}"
attributes:
reason: "{{ trigger.json.pause_reason }}"
since: "{{ trigger.json.paused_since }}"
Then set
QUEUE_WEBHOOK_URL=http://<home-assistant>:8123/api/webhook/scan-splitter-queue-CHANGE-ME.
Spending limit and paused processing
Mistral can refuse an account, for example because its spending limit is
reached, its quota is used up or its API key is rejected. In that case,
processing pauses instead of moving file after file to failed/.
A refusal is an HTTP 401, 402 or 403, or a 429 that is not the ordinary per-second rate limit. Mistral does not document how a reached spending limit is answered, so this errs on the side of pausing.
While paused:
- Files stay in their inboxes, including the one that hit the limit.
- The log shows
processing pausedwith Mistral's error text, and the webhook reports"paused": true. - Every
PAUSE_RETRY_MINUTES, one file is tried as a probe. Once it goes through, processing resumes on its own and the log showsprocessing resumed. That happens, for example, after you raise the limit or a new billing month starts.
Choosing the text source
Splitting and naming read the text of each page. Per input, it can come from
Mistral OCR (mistral) or from the Tesseract text layer (tesseract). The
Tesseract layer is added to every output PDF either way. By default, stacks
use mistral and scanner files use tesseract.
mistral |
tesseract |
|
|---|---|---|
| Cost | about $1 per 500 pages (batch) | none |
| Extra waiting time | about a minute per batch job | none |
| Tables, headers, footers | structured; headers and footers separate | plain text |
| Poor scans, handwriting | expected to be better (not measured) | expected to be weaker |
A comparison on one test stack used the same text layer for both sources. The stack held 80 real documents (156 pages of mostly clean office and medical scans):
mistral |
tesseract |
|
|---|---|---|
| Boundaries found | 77 of 80 | 75 of 80 |
| Clear misses | 0 | 2, both flagged for review (image-heavy pages with little text) |
| Same document type in the title | — | 64 of 81 documents |
| Same date | — | 73 of 81 documents; the differences favoured neither source |
The remaining misses of both runs were debatable cases. One example is two X-ray views of the same examination that had been filed as two documents.
In short: Mistral OCR splits somewhat better. For naming alone, as with
scanner files, the difference was not measurable. That is why the defaults
are what they are. If your scanner files are handwritten or of poor quality,
SCANNER_TEXT_SOURCE=mistral may be worth the cost.
Paperless-ngx input
Set PAPERLESS_URL and PAPERLESS_TOKEN to enable a third inbox,
paperless/inbox/. It is meant for documents that Paperless-ngx should name
and tag itself, for example with an AI tagger such as
Zettelrobbe. For each file:
- ocrmypdf adds the Tesseract text layer, with the same settings as for the other inputs. By default, no paid service is involved; see Content from Mistral OCR for the option.
- The PDF is uploaded through
POST /api/documents/post_document/, keeping its file name and addingPAPERLESS_TAGS, if set. - The container follows Paperless's consumption task until a document has
been created. Only then does the original move to
archive/, and the work copy is deleted.
Content from Mistral OCR
By default, the content field Paperless shows and searches comes from the
Tesseract text layer. In that text, tables lose their structure. Set
PAPERLESS_TEXT_SOURCE=mistral, and Mistral OCR reads each file before the
upload, through the batch API. Right after Paperless confirms the new
document, its content is replaced with Mistral's Markdown text. A lab report
then reads:
Laborbefund vom 30.09.2026
| Parameter | Ergebnis | Einheit | Referenz |
| --- | --- | --- | --- |
| Leukozyten | 6,2 | /nl | 3,9 - 10,2 |
The same page in Tesseract's text reads Leukozyten 6,2 /nl 3,9 - 10,2, one
line per row, with no columns.
- The PDF's own text layer stays Tesseract's, because only that one has word positions for search and selection in the file.
- Cost: Mistral OCR at batch price, about $1 per 500 pages. No chat model is involved.
- An AI tagger such as Zettelrobbe should see the better text. The content is replaced a few seconds after the document appears, and taggers usually poll less often. If yours reacts instantly, it may read Tesseract's text first.
- If replacing the content fails, the document stays in Paperless with Tesseract's text. A warning is logged, and the file is not uploaded again.
- A Mistral pause (see Spending limit)
holds this input too, because it uses Mistral. With the default
tesseract, the input keeps uploading during a pause.
Failure handling:
- Paperless unreachable, or token rejected. The file stays in the inbox, and the input retries after 5 minutes. A file that is still being consumed is not uploaded a second time after a restart: the task id is stored.
- Paperless rejects the document, for example as a duplicate. The file
moves to
failed/with Paperless's message in the.error.txt. - The same original dropped in twice. The file moves to
failed/and names the Paperless document it already became. Paperless's own duplicate check cannot catch this, because the text layer makes every upload a slightly different file. The container therefore keeps a register of uploaded originals inwork/paperless/uploaded.json. Remove an entry there to upload that file again.
Paperless decides by itself whether to run its own OCR. With the default
PAPERLESS_OCR_MODE=auto, it keeps the text layer added here. With redo
or force, it replaces it.
Sub-folders in paperless/inbox/ are allowed and mirrored in archive/, but
Paperless itself doesn't see them. Use PAPERLESS_TAGS, or Paperless
workflows, to sort documents.
Tested with Paperless-ngx 3.2. The task format of version 2 is supported as well.
A second Paperless user
Paperless makes the uploading user the owner of a document. When several people share one Paperless instance, each should own the documents scanned for them; otherwise permissions that limit someone to their own documents don't work.
Set PAPERLESS_2_TOKEN to the API token of a second Paperless user, and a
second inbox, paperless-2/inbox/, uploads as that user. Point a second
scanner profile, or a second network folder, at it.
| Setting | Default | |
|---|---|---|
PAPERLESS_2_TOKEN |
enables the second input | |
PAPERLESS_2_URL |
PAPERLESS_URL |
set it to upload to a different instance |
PAPERLESS_2_DIR |
/data/paperless-2 |
|
PAPERLESS_2_TAGS |
tag ids for this input | |
PAPERLESS_2_TEXT_SOURCE |
tesseract |
as PAPERLESS_TEXT_SOURCE |
PAPERLESS_MAX_WAIT_MINUTES applies to both. The second input works exactly
like the first, with its own duplicate register in
work/paperless-2/uploaded.json: the same scan may go to both users. Every
call it makes, uploads and content replacement included, uses its own token.
Shared tags, correspondents and document types
Paperless gives every new tag, correspondent and document type an owner, the user who created it, and other users only see those that have no owner or are shared with them. On an instance with several users, what one of them, or an AI tagger, creates is therefore invisible to everyone else. Paperless has no setting that makes new ones ownerless.
Set PAPERLESS_SHARE_TAGS=true, PAPERLESS_SHARE_CORRESPONDENTS=true and
PAPERLESS_SHARE_DOCUMENT_TYPES=true, each on its own, and the container
takes care of it. Every PAPERLESS_SHARE_TAGS_MINUTES (default 1), it
removes the owner and any explicit permissions from every object of those
kinds that has an owner. Every user with the global permissions for that
kind may then see and change it.
- The token needs permission to change other users' objects, in practice a
superuser's. With a weaker one,
tags could not be shared(orcorrespondents …,document types …) is logged once with HTTP 403, and the other kinds and the uploads carry on. The same happens while Paperless is unreachable. - Shared correspondents show every user who writes to whom: a bank, an employer, a doctor. Switch them on only where all users may see that.
PAPERLESS_SHARE_TAGS_READONLY(tags only) takes comma-separated tag names, for exampleai-processed, a marker an AI tagger uses to find work. Those tags keep their owner, every other user may see them, and nobody else may change them. A read-only tag without an owner is left alone with a warning: give it one in Paperless.- Paperless doesn't allow two ownerless objects of a kind with the same
name, e.g. a correspondent that you and an AI tagger both created. Such
an object stays as it is, and the log names it once as
correspondents not sharedwith itsidand the ownerless one it clashes with (same_name_as); all others are shared. Merge the two in Paperless, and it is shared in the next round. - A check with nothing to do costs one list request per kind. The log says
how many were shared, with their ids (
tags shared,correspondents shared,document types shared), never their names. - Storage paths and custom fields keep their owners.
Rate limits and cost
Mistral limits requests per second and tokens per minute. The limits depend
on the model and the account tier, and they are not published. Look yours up
in Mistral's console under API › Limits and set MISTRAL_MAX_RPS a
little below the requests-per-second limit of your MISTRAL_LLM_MODEL. If a
request still runs into the limit, all workers pause together and retry.
Requests for scanner and Paperless files take the next free slot, ahead of
those for a stack.
Splitting needs one request per 6 pages, and naming one request per document. A 500-page stack holding about 200 documents takes roughly 280 requests, which is about 20 minutes at 0.25 requests per second. A scanner file needs a single request and, with the default text source, no OCR. Every answer is cached in the work directory.
At the prices published in October 2026, a 500-page stack costs about US$1.60: about $1.00 for batch OCR and about $0.60 for the chat model. Check Mistral's pricing for current figures.
Privacy
Every page is sent to Mistral's API, first for OCR, then for splitting and naming. Only use this for documents you are entitled to process that way. If the documents belong to someone else, get their consent first, especially for health or financial records.
- The work directories hold the full OCR text of every input file and a
searchable copy of it. They are deleted
WORK_RETENTION_DAYS(30) days after the file was processed; until then,rebuildcan re-cut it. Folders of failed files stay until you delete them. - Logs contain file names, page numbers and counts, never document text. The file names are there on purpose: the log has to show which file is being worked on. Keep this in mind before sharing a log, and redact the names.
- The Paperless input keeps no copy once a document is confirmed in Paperless. Only the register of uploaded originals remains: checksum, file name, document id and date.
- The queue webhook sends counts only.
- Languages beyond the built-in ones are downloaded from GitHub (or
TESSDATA_URL) at startup. The request names only the language file.
Unraid
A Docker template and step-by-step instructions are in
unraid/.
Development
There is no local Python setup to maintain: the image contains everything,
and the tests run inside it. The image ships without pip, so the test run
installs it first with ensurepip.
docker build -t scanbutler:dev .
docker run --rm --entrypoint sh -v "$PWD:/src" -w /src scanbutler:dev \
-c "python3 -m ensurepip >/dev/null && python3 -m pip install -q pytest && python3 -m pytest -q"
The tests replace Mistral with a fake and need no API key. A smoke test checks the built image itself (watcher, health check, OCR models, file owners), also without a key:
tests/smoke.sh scanbutler:dev
CI runs both together with pip-audit, shellcheck and gitleaks on every
push. Images are built only
for version tags (v*), for linux/amd64 and linux/arm64. Each image is
signed with cosign and ships an SBOM and provenance.
See CHANGELOG.md for the release history and SECURITY.md for reporting vulnerabilities.
Contributing
Bug reports, ideas and pull requests are welcome. Please never attach real documents, OCR text or unredacted file names to an issue. The log lines and settings the issue form asks for are almost always enough. If a problem only shows with one particular file, describe it (pages, scanner, what is special about it) instead of sharing it.
For pull requests: keep the tests passing (see Development), add a test for new behaviour, and add a line to the changelog.
Disclaimer
This is an independent project, not affiliated with or endorsed by Mistral AI, Paperless-ngx or any scanner manufacturer. Product names belong to their owners.
Splitting and naming are automated and can be wrong. Check review.md and
the results before relying on them, for example before discarding paper
originals. The software comes without warranty; see the license.
License
Categories
Related apps
Explore more like this
Explore allDetails
ghcr.io/tom-joad/scanbutler:latestRuntime arguments
- Network
bridge- Shell
sh- Privileged
- false
- Extra Params
--security-opt no-new-privileges --memory=4g
Template configuration
Large scans to split. inbox/, output/, archive/ and failed/ are created here.
- Target
- /data/stacks
- Default
- /mnt/user/Documents/scanbutler/stacks
Scanner input, one document per file. Point the scanner's network-folder target at inbox/ below this path.
- Target
- /data/scanner
- Default
- /mnt/user/Documents/scanbutler/scanner
Optional Paperless-ngx input (needs PAPERLESS_URL): files get the text layer and are uploaded to Paperless.
- Target
- /data/paperless
- Default
- /mnt/user/Documents/scanbutler/paperless
Optional second Paperless input (needs PAPERLESS_2_TOKEN): uploads as another Paperless user, so its documents belong to that user.
- Target
- /data/paperless-2
Per-stack OCR text, plan.json, review.md and temporary page images (several GB while a large stack is processed). Holds full document text: delete stacks you no longer need.
- Target
- /config
- Default
- /mnt/user/appdata/scanbutler
User that owns new files (nobody)
- Default
- 99
Group for new files (users)
- Default
- 100
002: new files are writable for the users group, e.g. over SMB
- Default
- 002
Mistral API key
Model for splitting and naming
- Default
- mistral-large-latest
Requests per second, slightly below your account's limit for MISTRAL_LLM_MODEL (Mistral console: API > Limits)
- Default
- 1
Language for document titles, e.g. German. Empty: the language of each document.
Used in the file name when a document has no date
- Default
- undated
Tesseract languages joined with +, e.g. deu+eng+fra. deu and eng are built in; others are downloaded once into the Config folder
- Default
- deu+eng
Parallel OCR pages. auto: as many as the container's memory allows (--memory in Extra Parameters), at most one per CPU the container may use. Scanner and Paperless files go ahead of stacks
- Default
- auto
Optional: Paperless-ngx base URL, e.g. http://NAS-IP:8000. Enables the Paperless input.
API token of the Paperless user that should own the uploaded documents
tesseract: Paperless content from the text layer (free). mistral: content replaced with Mistral OCR text, tables kept (about $1 per 500 pages).
- Default
- tesseract
Optional: comma-separated Paperless tag ids to add on upload, e.g. 3,7
Optional: API token of a second Paperless user. Enables the second input (Paperless 2 path); its uploads belong to that user
Optional: Paperless URL for the second input. Empty: same as PAPERLESS_URL
Optional: comma-separated tag ids to add on upload through the second input
true: remove the owner from every Paperless tag, every minute, so all users see the same tags. Needs a superuser token.
- Default
- false
true: remove the owner from every Paperless correspondent, so all users see them (also who writes to whom). Needs a superuser token.
- Default
- false
true: remove the owner from every Paperless document type, so all users see them. Needs a superuser token.
- Default
- false
Optional: comma-separated tag names that keep their owner and are only visible to other users, e.g. ai-processed
Optional: POST queue counts here on every change and every 5 minutes, e.g. a Home Assistant webhook (http://HA:8123/api/webhook/ID). See the project README for the sensor config.
Text for splitting and naming stacks: mistral (Mistral OCR, paid) or tesseract (free, from the text layer)
- Default
- mistral
Text for naming scanner files: tesseract (free, from the text layer) or mistral (Mistral OCR, paid; worth it for handwriting or poor scans)
- Default
- tesseract
batch: half price, minutes per job. direct: immediate, full price.
- Default
- batch
Days after which a processed file's work folder in Config is deleted (rebuild works until then). 0 keeps everything
- Default
- 30
{title} is required, {date} is YYYY-MM-DD
- Default
- {title} {date}