Document intake for client-facing firms

Files what it can prove,
refuses what it can't.

An incoming attachment is identified from its filename, its sender, and its subject line -- three signals, tried in order of how much each can be trusted. No attachment is opened to work out what it is. What none of the three signals settle goes to a person, with the reason attached.

Never opens a file Every decision comes from the filename, the sender, or the subject
Three signals, ranked Filename first, sender only on an exact address, subject last
Refuses to guess Unknown, ambiguous, and archived attachments go to review, not a folder
Proven idempotent The same email run twice wrote zero duplicate rows (exec 51513 โ†’ 51514)

The blind spot

A file in the wrong folder looks filed. An open one is a privacy cost.

Most intake automation either drops what it cannot recognise, or opens the document to work it out. Both trade a visible refusal for an invisible mistake.

A guess dressed as a decision

A router that must always answer will file a numeric, opaque filename somewhere plausible. It lands in the wrong client's case, and nothing anywhere reports that it was ever uncertain.

Opening it to find out is its own cost

These attachments are passports, payslips, and bank statements. A router that reads the page to classify it has already sent a page of someone's ID somewhere -- whether or not the classification lands right.

A silent drop is a lost document

An unknown sender is not a reason to guess which client's case the attachment belongs to. It is a reason to file nothing, and say so -- filing a stranger's document into a real client's folder is the one mistake that is hard to undo.

How it works

Four steps, each one able to say "I don't know"

Every step either produces a real answer or an honest refusal. Nothing in between -- there is no step that opens a document to settle a tie.

  1. 01

    Read the signals

    The filename is tried first -- it is generated by a system and most specific. An exact sender address is tried next. The subject line, written by a human and prone to drift, is tried last.

    src/intake/identify.py
  2. 02

    Route

    Three lists come out: filed, review, and missing. Each is a different action -- file it, look at it, chase it -- never a single score that leaves the next step to be re-derived by hand.

    src/intake/process.py
  3. 03

    Read the page only on request

    Text is read only when the filename is silent and the calling event explicitly asks for it. No adapter this product ships sets that flag -- so this product never opens a file, by construction, not by a setting someone could leave on.

    classify_by_text.js ocr_jobs.js
  4. 04

    Record once

    One board row per attachment, keyed by a hash of its real contents. The same email, sent again, produces the same key -- so a re-run writes zero new rows instead of a duplicate.

    board.item_id()

Try it

Same filename, two routers, different outcome

Pick a real filename, shaped after the ones this router has actually been run against. Compare what a router forced to always answer would do against what this one actually does.

inbound attachment ยท one file at a time

Forced to always answer A router with no "I don't know"

Passport

The name states what it is. Filed as a passport.

confidence not applicable -- filename matched

Outcome Filed

Doc Intake Router Filename, then sender, then subject -- in that order

Passport

Same result -- the filename genuinely says what it is. Signal 1 handled it for free.

identify() -> passport | "filename contains a passport keyword"

Outcome Filed

What the real pipeline does with the outcome above

Filed Recognised documents, named with the case ref and date first
Review Everything unresolved, with the reason AND the fix attached
Missing Checklist items that have never arrived, in any email

All four classifications above were run for real against src/intake/identify.py on 2026-09-21. The sender-signal case uses the real GCash address logged in SENDER_RULES; the naive column is a category this router argues against, not a real competing product.

The rules

Three rules a non-programmer can read and argue with

Every keyword lives in one table for exactly this reason -- so a wrong rule is a one-line fix, not a code review.

Filename, sender, subject -- in that order

The filename is generated by a system and stays put. The sender address is fixed. A subject line is written by a human and drifts -- so it is tried last, and only after the other two abstain.

Ambiguity is a decision, not a tie-break

A filename matching two document types goes to review. Picking the higher-scoring match would file a passport as a payslip roughly as often as it got it right, silently.

An archive is not an unrecognised file

A .zip is caught before any keyword is even checked, and gets its own review reason -- "extract it" -- because that action is different from "classify it by hand."

identify() -- eight real filenames, one run

2026-09-21
identify() output for eight filenames, run 2026-09-21
Filename / signal Type
passport_copy_delacruz.pdf passport
enc-0995890_...pdf (sender: GCash) bank_statement
Payment_Advice.pdf (sender: HSBC) payment_advice
716299_Order_of_Payment_...pdf order_of_payment
receipt_..._cr_00219.zip archive -- never opened
bank_statement_or_payslip_aug.pdf ambiguous -- 2 types matched
scan0001.pdf unidentified -- left open

Four filed outright, one archive routed to "extract it," one ambiguous case routed to "choose which type," one left genuinely unidentified. None were opened to make that call.

What you get

Four things, in plain language

Written to be read by whoever owns the inbox, not only by whoever built the pipeline.

Filed folder

Recognised documents, named with the case reference and date first, so a folder listing is chronological without anyone setting a sort order.

Answers: what actually arrived, and where does it live?

Review queue with a reason

Three different refusal reasons -- unrecognised, ambiguous, or an unopened archive -- each carrying the specific next action, not just "needs review."

Answers: what does a human need to look at, and why?

Missing-checklist status

Computed from what has ever arrived across every email for that case, never just the one just processed -- so the chase never resets to the full list by accident.

Answers: what does this client still owe?

A clean "I don't know"

Every attachment none of the three signals could settle is counted, never silently dropped and never forced into the nearest category.

Answers: what does nobody actually know yet?

This router runs as a live n8n workflow. It has never seen a real inbox.

7 of 7 real-mail rows classified identically to the Python core, and a re-run of the same input wrote zero duplicate board rows -- both measured, not claimed. It sits on a real OCI n8n instance, not activated, because the Gmail adapter has no OAuth credential yet. This page says so on purpose.