Delivered, not sent
Alerts that prove they arrived
A malformed credential once made every alert fail silently while the job kept reporting success — the monitoring was down, and the monitoring said fine. The run now fails outright if an alert was raised and none was delivered. It's the most valuable gate in the codebase.
Evidence from 01 · Tacet
19 became 11
Every run said success
Matching on a phone number when an email was missing merged colleagues who share an office line: 19 people became 11 records, and nothing errored. What caught it was reading what each run actually did — created, updated, unchanged — not whether it finished. It now matches on the source’s own record id, and archives rather than deletes.
Evidence from 05 · Tandem
0 write tools
The model never holds the pen
The fixer’s model can read a project but cannot write to it. It returns a diff; code checks every path, commits it to its own branch, and runs the project’s tests at that exact commit. Nothing becomes a pull request until a person approves — and approving re-runs the checks first.
Evidence from 02 · Chief of Staff
55 of 60 wrong
First contact with real data
Its first run against a real ledger raised 60 findings, and 55 were wrong — tax-inclusive invoices, supplier bills carrying the supplier’s own reference, voided documents no rule should have read. Each one became a test. Rules written from outside a business break on first contact with its data.
Evidence from 03 · Tieout
Read back
“Succeeded” is not “changed”
Salesforce’s deploy tool twice reported success while changing nothing — once on a sharing default, once on a connection setting. Both showed up only when the setting was read back from the org. Every deploy is now retrieved and compared with what was meant, never taken on the deploy log’s word.
Evidence from 04 · Salesforce CRM Suite
Cycles, not status
Liveness you can't fake
A restart reported Running, health checks answered OK, and the status page showed no errors — while the old build held the port and the new one crash-looped every five seconds. The only tell was a counter that hadn't reset. I monitor work advancing, not processes existing.
Evidence from 01 · Tacet
5 posts, 1 row
Tested at the same time, not one at a time
A duplicate guard that checked first and wrote second passed a test that posted twice in a row — then created duplicate clients the moment a queue drained in parallel. The write itself is now the lock: five simultaneous posts produce exactly one row. A one-at-a-time test passes the broken version too.
Evidence from Client onboarding pipeline · More work
8.20, passed
The grader gets tested too
A grader that searched the whole reply for the expected answer found it in the citation — so a test set wrong about every document scored 8.20 against a pass mark of 8.0. The leak only ever flattered wrong answers, which is why normal runs never showed it. Measurements now get the same suspicion as the system they measure.
Evidence from Docket · More work
Never empty
A failed source is not an empty one
A cold-starting host answered 503, the fetch swallowed it, and a report rendered clean with an entire estate missing and nothing on the page saying so. Failures now surface as the report's top finding — this report is incomplete — and a failed pull no longer counts as coverage.
Evidence from 01 · Tacet · client audit
0.72 → 0.50
A guessed threshold fails quietly
The planned confidence line of 0.72 sat above every score the agent actually produced. It would have escalated every question while looking like a flawless refusal gate. Measured on a test corpus, answerable and unanswerable questions separate at 0.50. A threshold nobody measured is not cautious — it is wrong in a direction that happens to be quiet.
Evidence from Attest · More work
4/4 held
Tested on its failure paths
Off-topic, empty, partial, and citation-bait inputs — tested before the agent went anywhere near a user, and again on its current model through the live site. The re-run caught one gap: a message of only spaces got an answer. It is now rejected, and all four hold. Most agents are only ever tested on the questions they're expected to get.
Evidence from Source-cited support agent · More work
Sources returned
Answers you can audit
Each answer comes back with the source documents it was drawn from, so a wrong answer is traceable in seconds rather than argued about. On the current model, 7 of 7 answered questions carried their source. When the model lists its sources but leaves the line out, code after the model adds it back.
Evidence from Source-cited support agent · More work
Human gate
It escalates instead of guessing
Confident answers get learned behind a confidence gate; anything below it goes to a person. And when source documents change, the vector store changes with them — so the system doesn't quietly drift away from the truth.
Evidence from Source-cited support agent · More work