John Anastacio
0 — Portfolio
Open to AI Solutions Engineer roles · select projects — Remote worldwide

Automation that runs itself. And tells you when it doesn't.

John Anastacio — AI Solutions Engineer. I build AI agents and automations that run real operations, then instrument them so you're not the last to know when one stops earning: source-cited answers, tested failure paths, and an alert the moment something goes quiet.

Scroll

John Anastacio is an AI Solutions Engineer based in Bulacan, Central Luzon, Philippines, open to full-time AI Solutions Engineer and Solutions Architect roles and select projects, remote worldwide. He builds RAG pipelines, AI agents, self-running CRM systems, and end-to-end workflow automation with a focus on data integrity: source-cited retrieval, failure-path testing, confidence-gated learning, and vector stores kept in sync with source documents. Stack: n8n, GoHighLevel, Zapier, OpenAI, Anthropic Claude, Claude Skills, Google Vertex AI, Gemini, RAG, vector databases, Pinecone, Supabase, API and webhook integrations, plus Python and FastAPI backends and observability tooling shipped on Oracle Cloud, AWS (Lambda, DynamoDB, Terraform) and Azure Container Apps with pytest suites, GitHub Actions CI and pre-push regression gates. Background in enterprise IT infrastructure, presales solutions engineering, and systems administration.

0
automated tests passing across the five featured systems — re-run and counted, never recalled.
0
years debugging infrastructure and scoping impossible projects before automating them.
0%
success rate across 23,397 monitored runs on my own live automation estate — as measured 24 Aug 2026.
Selected Work

Curated
automations.

Selected projects and proof of concepts. Real systems, real results. Three of them are running right now →

01

Tacet — Automation Observability & Estate Audit

Telemetry that measures a live automation estate: n8n execution history, agent and backend events, dashboards, and quiet/failing alerts on an hourly run that aborts before ingest if its own test suite fails. LLM spend is priced per model, and anything it cannot price is reported as unpriced, never zero. The same codebase audits a client’s automation platform from evidence rather than interviews — what runs, whether it works, what it costs, where credentials are exposed — and names what it could not check. A second alarm runs on separate cloud infrastructure, because the machine hosting the first sleeps part of every day.

It watches the watchera run fails if an alert was raised and never delivered
785Tests, run before every ingest
HourlyAutomated collection
45 minFreshness threshold on each live runner
Pythonn8n APIGitHub ActionsAWS LambdaSQLite · DatasetteTelegram
Full case study →
02

Chief of Staff — Fixes That Wait for a Yes

A daily agent that reads six projects, runs shared read-only checks plus an adapter per project, and sends one severity-grouped digest to Telegram. Where a finding is fixable, a model proposes the fix — but it holds no write tools. It returns a diff, and code decides whether that diff may land: every path is checked (inside the project, never tests, never .git, no symlinks), the change is committed to its own branch, and the project’s own tests must pass at that exact commit. A pull request opens only after a person approves, and approving re-runs the verification on that commit first. Code never grades its own findings — each one is logged for a human verdict.

The model writes a diff. Code decides if it lands.and a person approves the exact commit
372Tests passing · 8 more need Docker
0Write tools the model holds
1 SHAWhat an approval is bound to
PythonClaude Code CLIGitSQLiteTelegram
03

Tieout — Deterministic Ledger Checker

Fourteen written rules run over a live Xero ledger every week without anyone starting it. In audit language it performs recomputation — each invoice recalculated from its own parts rather than its stated total taken on trust — item analysis across invoices and contacts, and four fraud checks aimed at where the money actually goes: two suppliers paid into one bank account, an account held in someone else’s name, details that changed since last month’s ledger, and a supplier paid into one of the firm’s own staff accounts. A rule that cannot run says why it could not — no prior ledger to compare against, no staff account list supplied — instead of passing quietly. There is deliberately no model anywhere in it: every check turned out to be a comparison or a total, so a model would only have made it slower, unreproducible, and moved client financial data out of the tenancy. Deterministic by construction — the report date is an input rather than a clock, money is never a float, and every run fingerprints both the data it read and the findings it produced, so two people on the same file can prove they ran the same thing. Exactly two files may open a network connection, and a test fails the build if a third ever appears.

It knows when it didn’t runthe watcher judges the report’s age, not its presence
286Tests, all passing
14Written rules
WeeklyUnattended, on AWS Lambda
PythonXero APIAWS LambdaClickUpn8n
04

Salesforce CRM Suite — Read Back, Not Trusted

Three connected builds on one real Salesforce org, one per layer. Declarative: new fields, a child object with a roll-up, and a Flow that creates the post-sale milestones and a follow-up task the moment a deal is Closed Won, with sharing locked to Private plus a rule for high-risk renewals. Apex: a trigger that logs data-quality problems as records are written, and a nightly batch job for the duplicates a trigger cannot see because the matching record was already saved clean. Integration: an inbound REST endpoint, an outbound callout to n8n proven by a real round trip, and a panel on the deal page. The lesson that carries: the deploy tool twice reported success while changing nothing, so every deploy is read back from the org and compared with what was meant.

“Succeeded” is not “changed”every deploy is read back from the org and compared with intent
27Apex tests, all passing
14Failures logged, each with its error and fix
2-wayREST in and out, proven live
SalesforceApexFlowLightning Web ComponentsNamed Credentialsn8n
05

Tandem — A GoHighLevel Sync That Never Deletes

A reusable n8n template that keeps GoHighLevel consistent with the system a business already runs on, so nobody logs in to add or remove records. One shared core owns every decision about GHL; each source plugs in through a thin adapter, and Salesforce contacts and deals are live today on a sandbox location. What it refuses is the product: a record deleted at the source is archived in GHL, never deleted; a deal with several possible contacts gets a labelled placeholder rather than a guess; an unchanged record is never re-sent; one bad record never stops a run. A 15-minute poll is the safety net, and a Closed Won in Salesforce pushes through a webhook guarded so it can never collide with a scheduled run.

What it refuses is the productarchive, never delete · a placeholder, never a guess
7 sClosed Won in Salesforce to GHL, measured
0Records it deletes
15 minPoll safety net under the push
n8nGoHighLevel APISalesforceWebhooksn8n Data Tables

More work

Built and tested, one line each. Open any row for the full record — what it does, what went wrong first, and the numbers behind it.

  • Tally — billing ledgerAWS Lambda · DynamoDB · Terraform · Stripe

    Records every Stripe payment exactly once, and reconciles both ways — its exhibit is a payment it never received and correctly admits to.

  • Source-cited support agentn8n · Cloudflare Workers AI · Pinecone · Google Drive

    Answers from its documents or escalates to a person; a new answer is learned only above a confidence gate, and vectors follow the source files as they change.

  • Attest — knowledge agentFastAPI · Pinecone · Azure OpenAI

    Refuses rather than guesses, and never shows “delivered” for a handoff it cannot see — a third, amber state instead.

  • Docket — document classifierPython · Azure OpenAI · CI

    Rules place what they can and a model sees only the rest: 300 documents, 169 by rules, 87 by the model, 44 left unplaced rather than guessed.

  • Evals regression gatePython · GitHub Actions · LLM judge

    A pull request that makes the answers worse cannot merge — with no silent fallback from the paid judge to the free one.

  • Parity — cross-cloud archiveAWS · Azure · Terraform · Salesforce

    The same data landed in two clouds and recounted independently on each side: 161 records in each, no mismatches.

  • Client onboarding pipelineMake · Notion · Google Drive · Calendly · Discord

    Form to CRM row to client folder to Discord to email in one flow; five simultaneous submissions produce exactly one client row.

  • Credential vaultPython · age encryption

    A key nobody claimed stops the import instead of being skipped, and no command prints a secret.

  • Doc Intake RouterPython · n8n

    Files documents from the filename and the sender’s record alone — no attachment is ever opened. Phase 1: the deterministic core.

  • Iris — iPhone as a webcamC# · Windows virtual camera · MIT

    Turns an iPhone into a Windows webcam over Wi-Fi with nothing installed on the phone. Public source, released.

Live · real Stripe sandbox and AWS account

Tally — Billing Event Ledger

Stripe payment events land in an AWS Lambda and are written once — and only once — into a DynamoDB ledger, with the whole stack provisioned from Terraform and a React console over the top. A duplicate webhook is refused by a conditional write rather than by a lookup that races, so two copies of the same event arriving together cannot both win. Reconciliation compares what Stripe sent with what was stored and reports both directions: sent but never stored, and stored but never sent.

What it admits

One ledger row is missing on purpose. It is the exhibit — a payment the ledger never received and correctly reports — not a bug to fix. Five defects appeared only once it was pointed at real accounts rather than fixtures. A “press to see the duplicate blocked” button was designed and then dropped: every press would have added a permanent phantom row, so the demo would have broken the guarantee it was selling.

2-wayReconciliation: missing and unexpected
5Defects found only against real accounts
0Rows the public page can write
AWS LambdaDynamoDBTerraformStripeReact
Running · demo corpus, open to try

Source-Cited Support Agent

An n8n agent over two Pinecone retrievers — company documents, plus answers it has learned — that answers from its sources or does not answer. A new answer is written back to the learned store only when its confidence clears 0.78; anything below that goes to a person on Slack. A Google Drive watch keeps the index in step with the source files: a changed file’s old vectors are deleted before the new ones go in, so the knowledge base cannot quietly drift from the documents it was built on.

What went wrong first

When the original cloud trial lapsed in August, answering stopped too — not just adding documents — because every question is embedded before the index can be searched. Swapping only the chat model would have produced a confident chatbot that could no longer see its own knowledge base, failing loudly nowhere. Answers and embeddings now both run on Cloudflare Workers AI.

0.78Confidence needed before an answer is learned
2Retrievers: company docs and learned answers
54Nodes in the live workflow
n8nCloudflare Workers AIPineconeGoogle DriveSlack
Built · the voice leg has never been spoken to

Attest — A Knowledge Agent That Refuses

A FastAPI service over Pinecone with Azure embeddings that answers only from a source document and carries the citation with the answer — and refuses outright rather than guess. When it hands a question to a person it confirms the handoff by reading the transport’s reply back. Its public endpoint reports an attempted handoff, never a delivered one, because the proof lives in a private log — so the console shows a third, amber state between delivered and failed.

What went wrong first

Six defects passed every unit test and appeared only against real services. Every citation was the literal string “blob” — a placeholder that satisfied “never answer without a source” while identifying nothing. The planned confidence threshold of 0.72 sat above every measured score: it would have escalated every question while looking like a flawless refusal gate. Measured, the line sits at 0.50. Nothing has been spoken aloud yet — the voice leg is configured, never exercised.

3Handoff states: delivered, failed, attempted
6Defects found only against real services
0.50Threshold, measured — the guess was 0.72
FastAPIPineconeAzure OpenAITelegramnginx
Built · CI green

Docket — Two-Tier Document Classification

Takes a pile of client documents, works out what each one is, decides what each client still owes, and drafts the chase letter. Rules go first — a keyword table over the filename: free, instant, correctable in one line. A model sees only what the rules could not place, in batches rather than one call per file, and it is allowed to return nothing, so “not decided” stays distinct from right and wrong. Asked for the real model without keys, it refuses rather than falling back to a stub.

What the measurements got wrong

Three measurement defects, all of which flattered the tester. A grader that searched the whole reply read the answer out of the citation — a corpus wrong about every document scored 8.20 against a pass mark of 8.0. Two categories described the same document, producing a 92.87% precision figure; with the answer key fixed it was one error, 99.80%, with no change to the system. The model tier’s 90.8% is self-agreement, not accuracy — its key came from the same model.

169 / 87 / 44Rules / model / left unplaced, of 300
6Model calls for 300 documents
8.20What a wholly wrong corpus scored — and passed
PythonAzure OpenAIGitHub Actions
Private · gates pull requests

Evals — The Regression Gate

Turns “did this answer get worse” from a judgement into a number CI can act on: a pull request that lowers the score cannot merge. Two thresholds rather than one — mean score and refusal rate — because a model that answers well but declines a third of the time is a regression a mean hides. Two backends: a free check of an answer’s shape and a paid judge of its quality, with no silent fallback between them. A check that could not run fails the run; it is never counted as a pass.

What it found, and what it cannot do yet

Pointed at Attest, it found two real defects, one of them a prompt injection the confidence gate answered. On one fix the shape score fell from 10.0 to 9.29 while quality rose from 7.0 to 9.43 — which is why the two are never averaged. And the judge’s own run-to-run noise measured +0.43 against a 0.5 drift tolerance, so it cannot yet gate on a single mean.

2Thresholds: mean score and refusal rate
+0.43Judge noise, against a 0.5 tolerance
2Real defects found in Attest
PythonGitHub ActionsLLM judge
Built · real AWS and Azure accounts

Parity — Cross-Cloud Reconciliation

An event pipeline lands the same batch in two independent clouds, then recomputes row count, duplicate keys and a content digest separately on each side rather than trusting either side’s “write succeeded”. Writes are create-only with byte-conflict checks, and the Salesforce write-back locks the row and checks its version inside one Apex transaction. A mismatch now lands in the Salesforce CRM Suite’s own audit trail instead of only a metric. It is an incremental archive of six fields, not a full CRM backup.

What went wrong first

The first cloud leg was already called deployed and had never processed a real record — two silent bugs, invisible to every test that ran on an empty batch, found only by reading its own live logs. Repairing it meant restoring 42 missing copies from the original bytes. A live end-to-end call into Salesforce then caught a field-length bug the mocked tests had missed.

161 / 161Records in each cloud, no mismatches
42Missing copies restored from original bytes
2Silent bugs found in live logs, not tests
AWS LambdaAzureTerraformSalesforcePython
Built on the live Make account, entirely through its API

Client Onboarding Pipeline

One Make scenario, 19 modules, two routes behind a router — because the account allows only two active scenarios. A form submission creates a Notion CRM row and a Google Drive client folder, writes the folder link back, posts to Discord and emails the lead. A Calendly booking or cancellation updates the CRM row, creates a call-prep document and posts to Discord. Every outward step was checked at its destination — messages read back from the channel, the email found in the inbox — never by a 200.

What went wrong first

The first duplicate guard asked “does this exist?” and then acted. It passed a test that posted twice in a row, and produced duplicate client rows the moment a 13-item queue drained, because Make drains queues in parallel. The fix makes the write itself the lock: an insert that fails if the key already exists. Five simultaneous posts now produce exactly one row — and a one-at-a-time test passes the broken version too, which is why the first one lied.

19Modules in one scenario
5 → 1Simultaneous posts to client rows created
2Routes: intake and bookings
MakeNotionGoogle DriveCalendlyDiscordGmail
In daily use · not yet a backup

Credential Vault

Every credential in the workspace, encrypted at rest, plus a tool that rebuilds any project’s environment file from it. Two separate keys open ordinary credentials and dangerous ones — anything that can move money or destroy infrastructure sits behind the second — and the master key that opens everything is deliberately not on the machine. The load-bearing behaviour is a refusal: if an environment file holds a key that no entry owns and no rule has excused, the import stops rather than proceeding. There is no command that prints a secret, by design.

What it is not

It is not yet a backup. It sits on the same disk as the files it protects, so a dead disk still costs everything. The README says so in bold, and an off-machine copy is the open item.

2Keys, split by blast radius
0Commands that print a secret
Off-boxWhere the master key lives
Pythonage encryptionEnvironment manifests
Phase 1 · the deterministic core

Doc Intake Router

Files documents without opening them. Every filing decision comes from the filename and the sender’s record; no attachment is read, embedded or placed in a model’s context — and no setting can change that, because the code to read one does not exist. Output is three lists, never a percentage: filed (with the exact folder and new name), review (with the reason), and missing (what to chase).

Four rules it will not break

An unknown sender files nothing. A filename matching two document types goes to review, not to the better guess. Nothing is dropped. An unrecognised extension goes to review, not the bin. The trap it guards against: a growing review pile invites loosening a rule, which turns visible uncertainty into invisible misfiling.

3Output lists, never a percentage
4Rules, each pinned by a test
7 / 7n8n version matching the Python; not switched on
Pythonn8n
Shipped · public, MIT · not AI work

Iris — An iPhone as a Windows Webcam

A Windows virtual-camera driver and client that turn an iPhone into a webcam over Wi-Fi, with nothing installed on the phone. Released as v1.1.0, a single executable. It passes frames across a Windows session boundary without locks, and its test suite was checked by mutation — the code deliberately broken to prove the tests fail when they should.

What to know

It is here because its source is public — anyone can read how it works. It is not AI work, and the binaries are unsigned.

v1.1.0Released
MITPublic source
0Apps installed on the phone
C#WindowsVirtual camera
Reliability

How you'll know
it actually works.

Delivered, not sent

Alerts that prove they arrived

A malformed credential once made every alert fail silently while the job kept reporting success — the monitoring was down, and the monitoring said fine. The run now fails outright if an alert was raised and none was delivered. It's the most valuable gate in the codebase.

Evidence from 01 · Tacet

19 became 11

Every run said success

Matching on a phone number when an email was missing merged colleagues who share an office line: 19 people became 11 records, and nothing errored. What caught it was reading what each run actually did — created, updated, unchanged — not whether it finished. It now matches on the source’s own record id, and archives rather than deletes.

Evidence from 05 · Tandem

0 write tools

The model never holds the pen

The fixer’s model can read a project but cannot write to it. It returns a diff; code checks every path, commits it to its own branch, and runs the project’s tests at that exact commit. Nothing becomes a pull request until a person approves — and approving re-runs the checks first.

Evidence from 02 · Chief of Staff

55 of 60 wrong

First contact with real data

Its first run against a real ledger raised 60 findings, and 55 were wrong — tax-inclusive invoices, supplier bills carrying the supplier’s own reference, voided documents no rule should have read. Each one became a test. Rules written from outside a business break on first contact with its data.

Evidence from 03 · Tieout

Read back

“Succeeded” is not “changed”

Salesforce’s deploy tool twice reported success while changing nothing — once on a sharing default, once on a connection setting. Both showed up only when the setting was read back from the org. Every deploy is now retrieved and compared with what was meant, never taken on the deploy log’s word.

Evidence from 04 · Salesforce CRM Suite

Cycles, not status

Liveness you can't fake

A restart reported Running, health checks answered OK, and the status page showed no errors — while the old build held the port and the new one crash-looped every five seconds. The only tell was a counter that hadn't reset. I monitor work advancing, not processes existing.

Evidence from 01 · Tacet

5 posts, 1 row

Tested at the same time, not one at a time

A duplicate guard that checked first and wrote second passed a test that posted twice in a row — then created duplicate clients the moment a queue drained in parallel. The write itself is now the lock: five simultaneous posts produce exactly one row. A one-at-a-time test passes the broken version too.

Evidence from Client onboarding pipeline · More work

8.20, passed

The grader gets tested too

A grader that searched the whole reply for the expected answer found it in the citation — so a test set wrong about every document scored 8.20 against a pass mark of 8.0. The leak only ever flattered wrong answers, which is why normal runs never showed it. Measurements now get the same suspicion as the system they measure.

Evidence from Docket · More work

Never empty

A failed source is not an empty one

A cold-starting host answered 503, the fetch swallowed it, and a report rendered clean with an entire estate missing and nothing on the page saying so. Failures now surface as the report's top finding — this report is incomplete — and a failed pull no longer counts as coverage.

Evidence from 01 · Tacet · client audit

0.72 → 0.50

A guessed threshold fails quietly

The planned confidence line of 0.72 sat above every score the agent actually produced. It would have escalated every question while looking like a flawless refusal gate. Measured on a test corpus, answerable and unanswerable questions separate at 0.50. A threshold nobody measured is not cautious — it is wrong in a direction that happens to be quiet.

Evidence from Attest · More work

4/4 held

Tested on its failure paths

Off-topic, empty, partial, and citation-bait inputs — tested before the agent went anywhere near a user, and again on its current model through the live site. The re-run caught one gap: a message of only spaces got an answer. It is now rejected, and all four hold. Most agents are only ever tested on the questions they're expected to get.

Evidence from Source-cited support agent · More work

Sources returned

Answers you can audit

Each answer comes back with the source documents it was drawn from, so a wrong answer is traceable in seconds rather than argued about. On the current model, 7 of 7 answered questions carried their source. When the model lists its sources but leaves the line out, code after the model adds it back.

Evidence from Source-cited support agent · More work

Human gate

It escalates instead of guessing

Confident answers get learned behind a confidence gate; anything below it goes to a person. And when source documents change, the vector store changes with them — so the system doesn't quietly drift away from the truth.

Evidence from Source-cited support agent · More work

What stays human

No agent I build sends a customer-facing message on its own, commits money, or resolves the edge cases it was never confident about. Those escalate — by design, not by accident. A scope that admits its limits is the only kind worth signing off on.

About

John Anastacio builds self-running systems that turn underused tech stacks into ROI-generating operations.

AI Solutions Engineer with almost a decade of experience in enterprise IT infrastructure and presales solutions engineering. I design and build RAG pipelines, AI agents, self-running CRM systems, and end-to-end workflow automation — for real estate agents, financial advisors, SaaS founders, and consultants.

My builds treat data integrity as a feature: source-cited retrieval, failure-path testing, and knowledge bases that stay in sync with their source documents — reliable, documented systems that keep working long after launch.

More recently that work has turned outward: telemetry platforms that measure whether an automation estate is actually running, and audits that report what a business has, what it costs, and what has quietly stopped working — stating plainly what was measured and what wasn't.

AI & agents

  • OpenAI · Claude · Gemini (Vertex AI)
  • Claude Skills & agent orchestration
  • Prompt engineering & failure-path testing
  • n8n, Zapier & Make

Data & RAG pipelines

  • Pinecone & vector stores
  • Supabase · Postgres · DuckDB · SQLite
  • Embedding & ingest pipelines (CRUD sync)
  • Source-cited retrieval

Observability & audits

  • Run & execution telemetry
  • Health, freshness & liveness monitoring
  • Failure alerting with delivery verification
  • LLM token & spend attribution
  • Automation estate audits & credential-exposure scanning

Backends & APIs

  • Python · FastAPI services
  • REST & webhook integrations
  • Scheduled jobs & background workers
  • Telegram & Slack bot interfaces

Ship & operate

  • AWS Lambda · DynamoDB · Terraform
  • Azure Container Apps jobs · Azure OpenAI
  • Oracle Cloud · systemd · Windows Task Scheduler
  • SSH deploys with post-deploy verification
  • pytest suites, GitHub Actions CI & pre-push regression gates
  • Encrypted credential vault · least-privilege job identities
  • Health checks, alerting & runbooks

CRM & workflow

  • GoHighLevel
  • Pipeline design
  • Lead capture & follow-up

Enterprise infrastructure

  • Fortinet · Palo Alto · HPE
  • NAS / SAN · VMware · Windows Server
  • SOPs, runbooks & change management
  • SLA management
Mar 2026 — Present
Independent AI Systems Engineer
Independent practice · Remote
Building and operating AI and automation systems as production services — observability and cost accounting for live automation estates, client estate audits, and evidence-gated delivery where a change has to prove itself before it ships.
Feb 2018 — Mar 2026
Presales Solutions Engineer
Concentrix · Philippines
Scoped and recommended network, server, and storage architectures across security, compute, and storage — from targeted upgrades to full-scale migrations for multiple clients.
Apr 2016 — Jun 2017
Windows System Administrator
Hewlett Packard Enterprise · Taguig
Managed enterprise IT operations across incident, change, and project management within strict SLAs in a multinational data center — authoring change requests and runbooks to standardize procedures.
Sep 2015 — Dec 2015
RIM Bootcamp — System Administration Training
Fujitsu Philippines Global Delivery Center
2010 — 2015
B.A.Sc. — Electronics & Comms Engineering
Polytechnic University of the Philippines
Methodology

How the work
actually gets done.

01

Discovery & Scoping

We find where your team is losing the most time and agree on exactly what success looks like before anything gets built.

02

Architecture Design

You get a clear blueprint of how the automation works, what it connects to, and how it handles edge cases — no surprises mid-build.

03

Build in Slices

One rule decides the pace: the next slice isn't built until the current one has run on real data and been watched doing it. Automations don't fail on the inputs anyone thought about up front, and a small live slice is the cheapest place to meet the ones nobody could. You see working pieces early, and there is never a "we built it all and now nothing works" moment.

04

Prove It Works

Before anything goes live, we agree on what "working" means in numbers — accuracy on your own real cases, what escalates to a human, and what happens when it's wrong. Then it gets tested against that, not against a demo.

05

Deploy & Sustain

You get a live system with full documentation — and an optional retainer for the work that keeps an agent honest after launch: monitoring, re-testing when models change, and tuning as your process moves.

Then every change after that clears seven gates.

The five steps above run once. These run every time anything changes afterwards — a tweak, a fix, a new rule — because that is when working systems quietly stop working. A change clears all seven or it doesn't ship.

Gate 01
Say it as a claim

"Changing this should produce more of that" — never "make it better." A vague suggestion can't be tested, so it isn't accepted.

Gate 02
Check history first

We keep a written registry of what has already been tried and disproven. If it's in there, we stop and point at the earlier result. This step saves the most time of any.

Gate 03
Reproduce the control

Run the unchanged system first and confirm it still gives the documented numbers. You can't measure a change against a ruler you haven't checked.

Gate 04
One variable

One change per experiment, across several conditions rather than the one convenient case — so when the result moves, we know what moved it.

Gate 05
Promote, reject, or shelve

It has to beat the control where it matters, not on a lucky average. A win that rests on one good slice is treated with suspicion, not celebrated.

Gate 06
Review the actual diff

Anything touching money, enforcement, or alerts is reviewed against the real change before it goes anywhere.

Gate 07
Record it either way

It ships through one verified deploy, or the dead end is written down. A rejection isn't a wasted experiment — it permanently narrows the search.

Underneath
A net that fails closed

The gates prove a change helps. An automatic regression check proves it broke nothing else, and it blocks rather than warns. A change passes both.

The whole method, written down.

The seven gates above, the build loop behind step 03, and how the same discipline gets pointed at a problem in your environment — written out in full, with what gets recorded when something is rejected. Worth reading before you hire anyone, including me.

How We Work
Offers

Tailored
AI solutions.

Start here

Agent Opportunity Audit

Before anyone writes a line of agent code, we find out which of your processes is actually worth automating — and which would quietly eat a budget and ship nothing. Up to three processes, measured out of your own systems rather than estimated, and scored on volume, error cost, data readiness and how much of the judgement can't be handed over. You get a ranked build queue and a risk register, walked through live. Sometimes the answer is "don't build this yet" — which is worth knowing before you spend on a build.

Customone week · credited toward the build

Then the queue decides which of these it becomes —

Workflow Automation

One repetitive, time-consuming workflow — scoped, built, documented, and handed over. Best for small teams who want a specific bottleneck gone without a long agency engagement.

$2,500per project *
Most popular

AI Agent Build

An agent that runs a real process end to end — retrieval, decisions, escalation — with an accuracy target you sign off on before launch and a human gate on the calls it shouldn't make alone.

$5,000per project *

AI Consultancy & Scale

For teams running complex, multi-stage operations, or who need someone to work out what's worth automating before anyone builds. Scope and pricing shaped around the engagement — including ongoing agent operations once systems are live.

Customscope-dependent

* fixed price per project, with an optional monthly retainer for monitoring, re-testing, and tuning after launch · all engagements remote-first, worldwide

Built for your bottleneck,
not off-the-shelf.

Works with founders, ops leaders, and product teams who take automation seriously — and want it built once, built right. Currently open to full-time AI Solutions Engineer / Architect roles alongside select projects.

Not sure it's a fit? Answer 10 questions first →
Bulacan, Philippines · Available worldwide
SuMoTuWeThFrSa
· 30-min Free Workflow Audit