A guess dressed as a decision
A router that must always answer will file a numeric, opaque filename somewhere plausible. It lands in the wrong client's case, and nothing anywhere reports that it was ever uncertain.
Document intake for client-facing firms
An incoming attachment is identified from its filename, its sender, and its subject line -- three signals, tried in order of how much each can be trusted. No attachment is opened to work out what it is. What none of the three signals settle goes to a person, with the reason attached.
The blind spot
Most intake automation either drops what it cannot recognise, or opens the document to work it out. Both trade a visible refusal for an invisible mistake.
A router that must always answer will file a numeric, opaque filename somewhere plausible. It lands in the wrong client's case, and nothing anywhere reports that it was ever uncertain.
These attachments are passports, payslips, and bank statements. A router that reads the page to classify it has already sent a page of someone's ID somewhere -- whether or not the classification lands right.
An unknown sender is not a reason to guess which client's case the attachment belongs to. It is a reason to file nothing, and say so -- filing a stranger's document into a real client's folder is the one mistake that is hard to undo.
How it works
Every step either produces a real answer or an honest refusal. Nothing in between -- there is no step that opens a document to settle a tie.
The filename is tried first -- it is generated by a system and most specific. An exact sender address is tried next. The subject line, written by a human and prone to drift, is tried last.
src/intake/identify.py
Three lists come out: filed, review, and missing. Each is a different action -- file it, look at it, chase it -- never a single score that leaves the next step to be re-derived by hand.
src/intake/process.py
Text is read only when the filename is silent and the calling event explicitly asks for it. No adapter this product ships sets that flag -- so this product never opens a file, by construction, not by a setting someone could leave on.
classify_by_text.js
ocr_jobs.js
One board row per attachment, keyed by a hash of its real contents. The same email, sent again, produces the same key -- so a re-run writes zero new rows instead of a duplicate.
board.item_id()
Try it
Pick a real filename, shaped after the ones this router has actually been run against. Compare what a router forced to always answer would do against what this one actually does.
The name states what it is. Filed as a passport.
confidence not applicable -- filename matched
Same result -- the filename genuinely says what it is. Signal 1 handled it for free.
identify() -> passport | "filename contains a passport keyword"
What the real pipeline does with the outcome above
All four classifications above were run for real against
src/intake/identify.py on 2026-09-21. The sender-signal case
uses the real GCash address logged in SENDER_RULES; the naive
column is a category this router argues against, not a real competing
product.
The rules
Every keyword lives in one table for exactly this reason -- so a wrong rule is a one-line fix, not a code review.
The filename is generated by a system and stays put. The sender address is fixed. A subject line is written by a human and drifts -- so it is tried last, and only after the other two abstain.
A filename matching two document types goes to review. Picking the higher-scoring match would file a passport as a payslip roughly as often as it got it right, silently.
A .zip is caught before any keyword is even checked, and gets its own review reason -- "extract it" -- because that action is different from "classify it by hand."
| Filename / signal | Type |
|---|---|
| passport_copy_delacruz.pdf | passport |
| enc-0995890_...pdf (sender: GCash) | bank_statement |
| Payment_Advice.pdf (sender: HSBC) | payment_advice |
| 716299_Order_of_Payment_...pdf | order_of_payment |
| receipt_..._cr_00219.zip | archive -- never opened |
| bank_statement_or_payslip_aug.pdf | ambiguous -- 2 types matched |
| scan0001.pdf | unidentified -- left open |
Four filed outright, one archive routed to "extract it," one ambiguous case routed to "choose which type," one left genuinely unidentified. None were opened to make that call.
What you get
Written to be read by whoever owns the inbox, not only by whoever built the pipeline.
Recognised documents, named with the case reference and date first, so a folder listing is chronological without anyone setting a sort order.
Answers: what actually arrived, and where does it live?
Three different refusal reasons -- unrecognised, ambiguous, or an unopened archive -- each carrying the specific next action, not just "needs review."
Answers: what does a human need to look at, and why?
Computed from what has ever arrived across every email for that case, never just the one just processed -- so the chase never resets to the full list by accident.
Answers: what does this client still owe?
Every attachment none of the three signals could settle is counted, never silently dropped and never forced into the nearest category.
Answers: what does nobody actually know yet?
7 of 7 real-mail rows classified identically to the Python core, and a re-run of the same input wrote zero duplicate board rows -- both measured, not claimed. It sits on a real OCI n8n instance, not activated, because the Gmail adapter has no OAuth credential yet. This page says so on purpose.