Case studies

What they arrived with. What they left with.

Four clients and two products of my own, written up the same way: the before, the after, and the one call that made the difference. I was architect and sole developer on all six.

Client systems run in the clients’ own accounts, so their figures are theirs to publish. The only numbers on this page are mine.

SSynthetic & Human to Ush

Version 16 is running.

Built this week
Slots learn from which ones get accepted and declined, both sides
Still working from week 12
Screenshot a WhatsApp thread, get the slots that fit
If there is a week 17
The day on a map, with travel time between meetings
Book week 17Stop here
6systems, written up the same way
3countries: UK, Germany, Israel
40weeks the longest client bought, one at a time, back when the unit was a week
1Mpages a day through SignalsAPI, mine
0people on duty to keep it running

Six systems

Before and after, in one line each.

Client · UKExecutive assistants2025–2026

Booking a meeting went from a chain of emails to one exchange.

Arrived with

Screen mockups and a wish list. The judgement that picks a meeting time — habits, travel, who outranks whom — lived in one assistant’s head.

Left with

A scheduler that learns each boss’s hours from the calendar, ranks slots for least travel and waiting, improves from which slots get accepted, and answers a WhatsApp screenshot with the slots that fit.

The call

The spec was a workflow: fixed steps. Scheduling isn’t; it’s a ranking. So the screens they arrived with were never built. A new version every week, tried by real users, replaced the plan.

Tryable at week 4Useful at week 16Still buying at week 40
For the engineer

Hard constraints filter, weighted preferences rank, accept/decline history adjusts the weights. The model reads requests and writes the explanation; it never picks the slot. Flask, one worker, PostgreSQL 16 for sessions, queue and storage. Every query scoped to a user_id by construction. Gmail read and compose only; no send scope, on purpose.

app.ush.example/requests/1184
Coffee with J. Okafor · 30 minfrom WhatsApp screenshot
Any chance Thu or Fri for a quick coffee?
Let me check with her diary.
Thu 11:308 min from the 10:30 in Holborn · no gap after · both freeBest
Thu 15:00free, but 40 min waiting before the 16:00Good
Fri 09:00before her usual start · declined twice at this hourAvoid
Draft the replya person sends it
Thursdaytravel time between meetings
091011121314 Board prep09:00–09:50 · office R. Achebe10:30–11:15 · Holborn 8 min walk Coffee · J. Okafor11:30–12:00 · Holborn 14 min walk Lunch · investors12:30–13:30 · Mayfair learned from the diary: she starts at 09:00

Schematic — shape only

Forty weeks, bought one at a time

4 16 40
building tryable, week 4 useful for real work, week 16

Each square is a week they chose to buy, back when the unit was a week, and could have declined. Schematic — shape only

The only hard number it has

Before a chain, and on After Thu 11:30 works, 8 min from Holborn Confirmed. 1exchange,down from a chain a person still sends it

Schematic — shape only

Client · GermanyReal-estate deal management10 weeks

Back to adding features with Codex, without breaking the rest.

Arrived with

A repo vibe-coded in Codex, deployed, serving nobody. One CRM hardwired, half the credentials in the code, half the features working — and nobody, Codex included, could say which half or why.

Left with

The same product as software. Any customer signs up and sets their own process. Any CRM plugs in. It keeps running when a CRM’s API doesn’t match its own docs — in German real estate, often.

The call

The problem wasn’t the code, it was the shape: business rules in a spreadsheet, a production system reading them and guessing from a log. Two queues and adapters replaced that — the one shape where a vibe-coding tool can add a feature without touching anything else.

2 CRMs connectedNext CRM = 1 adapterCredentials in code: 0
“We came back to our trusted process. We ask Codex for another feature, and it just works.”
KClient, KIMA · investor, manager, domain expert
For the engineer

Adapters normalise vendor events into one shape in and vendor calls out of one shape out, so the rules engine never knows which CRM it’s talking to and a bad API can’t reach it. Every event and action carries a tenant; rules are tenant data. Adapters are written against observed behaviour, not documentation, and tested in isolation.

app.kima.example/deals
Deals · Müller ImmobilienCRM A connected tenant: mueller
New lead3
Goethestraße 12
owner called, 14:10 auto
Am Park 4
calling owner at 17:00 queued
Viewing5
Lindenweg 8
Tue 14:00 booked, 2 buyers auto
Lindenweg 8
your task: open the door
Documents → notary3
Bahnhofstraße 3
exposé filled and sent auto
Rosenweg 21
your task: go to the notary
events in: 1,204 today · actions out: 318a person gets a task only when a person is needed

Before — roughly half worked; nobody knew which half

each feature built on its own spreadsheet one CRM a log read directly, guessed at

Schematic — shape only

After — every feature through the same two queues

CRM A CRM B CRM n adapter rules enginetenant rules as data adapter eventsactions back into the same CRMs · every event carries a tenant

Schematic — shape only

SSynthetic & Human to KIMA

Version 10 is running.

Built this week
Second CRM connected. One adapter, one prompt, contained and tested.
Still working from week 6
Tenants register and define their own process
If there is a week 11
Your call. The code and the accounts are yours already.
Book week 11Stop here — they did
Client · IsraelKnowledge graph startup15 PhDs, no engineers

A two-day laptop run moved to the cloud, and one scientist’s tool now reaches all fifteen.

Arrived with

A well-built pipeline that only ran on laptops: a day or two per run, results visible to one person. Four scientists had vibe-coded the same graph browser without knowing it.

Left with

The pipeline in the cloud on compute that comes up for a run and shuts down after. One shared store, secured, team and user access kept apart. An API to the graph that answers what a claim rests on with the paper and page, or the memo, it came from, and a place to deploy their own tools.

The call

Don’t rewrite the pipeline. It was good; the problem was where it ran. Move it, build the platform around it, and make deploying a vibe-coded tool cheap enough that sharing beats rebuilding.

Laptop run: 1–2 daysSame browser built 4× → 1×2 access layers
For the engineer

Same stages, now against an object store and a compute layer that scales to zero. State lives in the store, not on a machine, so a run can be inspected by someone who didn’t start it. A query API over the graph is what every tool reads; the scientists’ apps deploy beside it without an engineer in the loop.

graph.internal.example/browse
Graph browserdeployed by Dana · used by 15
Which claims depend on MOF-5 stability?
MOF-5 stability Chen 2023, p.4internal memo 118

Schematic — shape only

Runsshared store
#212 · full pipelinedone
started by Yael · 4 nodes up → 0 after
#213 · re-embed batchrunning
started by Omer · 2 nodes up
#211 · full pipelinedone
started by Dana · inspected by 6 others
access team users · API only

Where a run lives

Before day 2 of 2 1 person sees it After up for the run, then off one shared store 15 people see it

Schematic — shape only

The same tool, built four times

Before 4 built · 1 user each nobody knew the others existed After deployed 1 built · 15 users next one: a prompt away

Schematic — shape only

Client · financed by VolkswagenJob-market data for HR planningMultinational HR department

Volkswagen’s HR insights, made real time.

Here as proof of the other kind of client: not an idea or an MVP, but an established company with a working system, a precise ask, and no wish to hear about a rewrite.

Arrived with

A system that tracked who was opening which jobs, which skills were in demand and which were being replaced by AI — for Volkswagen’s HR department to plan with. It worked, and it was slow as hell: once a month a batch job ran over the postings, took a few days, and produced a handful of reports that HR read for the following month. A vacancy opened on 16 September reached a reader around 5 October.

Left with

The same data, real time. A posting opened today is in the database today. Many years of postings answer a query in a blink, and instead of reading a report the team asks a question in plain language: the AI writes the SQL, the query runs, the data comes back.

The call

Leave the collection alone; they hadn’t asked for it and it wasn’t the problem. The month between a posting and a reader was. Four separate changes that are easy to blur into one: latency, storage, access, interface. The reports went away entirely — they see the data first and query it after.

Posting → insight: a month → same dayMonthly reports: 0, queries insteadQuestion → SQL → answer
For the engineer

Postings land in an AWS data lake. A DAG runs over the lake: extract, convert, pull the required skills out of each posting, build the analytics on top — that is where the month went. The analytics now land in ClickHouse, which is what makes a query across the whole history fast rather than a run across it. A model turns a plain-language question into SQL against ClickHouse; the SQL executes; the answer is the rows, never the model’s memory.

insights.hr.internal.example/ask
Ask the dataAI writes the SQL
Which skills have stopped appearing in new postings this quarter, and which ones took their place?
SELECT skill, count() AS postings
FROM job_postings
WHERE opened >= toStartOfQuarter(today())
  AND skill NOT IN (SELECT skill FROM job_postings
    WHERE opened BETWEEN today() - 372 AND today() - 280)
GROUP BY skill ORDER BY postings DESC LIMIT 5
agent workflow designnew this quarter41
LLM evaluationnew this quarter27
manual test scriptinggone since May0
whole history scannedClickHouse · 0.2 s
Freshnessreal time
Posting openedWed 16 Sep, 09:12
In the databaseWed 16 Sep, 09:14
Queryable by HRWed 16 Sep, 09:14
Old cycle: visibleMon 5 Oct
Reports produced this month: 0. Questions asked: 212.

Before — a posting waits for the monthly run

waits for themonthly run the run takesa few days HR readsthe reports 16 Sep 1 Oct 5 Oct Nov posted read a month, on average, from posting to reader once a month · discrete · batch

Schematic — shape only

After — the data first, the question second

data lake DAGsame day ClickHouseyears of postings HR asksplain language AI → SQLwrites the query runs the rows come back

Schematic — shape only

MineHiring signals, sold through one API

A million pages a day, and nobody on duty.

Here as evidence the method builds real things, not as a client outcome. These are the only figures on this site I am free to publish.

Started with

Hundreds of career pages to read every day, and nobody to run the machine. At that size a system that needs an operator is a system that stops.

Runs as

Thirty-odd services, each with one job. Two things read the results — a filter console and a public site — and neither can ask an expensive question by accident.

The call

Nothing reaches the web on its own. One fetch service — real browsers on rented addresses — and everything else asks it for a page. One bill to read, one place to break, no service quietly growing its own crawler.

For the engineer

Sources go to the fetch pool, then a crawler normalises and dedupes, then classification into a store the API reads from. The company graph stores claims about companies, not companies. The filter console splits counting from searching into routes that cannot express each other. The mail service is in the picture because I run it, not because anything here uses it.

api.signalsapi.example/v1/signals
Todayhealthy · 0 on duty
Pages fetched1,000,000
Sources read457
Services live30+
Browsers in the poolrented, one bill
Expensive queries by accident0 possible
GET /v1/signals?company=northwind&since=7d 200 OK · 41 ms { "company": "northwind", "signals": [ { "role": "Staff Engineer, Payments", "seen": "2026-09-14", "source": "careers page" }, { "role": "Head of Risk", "seen": "2026-09-12", "source": "ATS board" } ], "count": 2, "pages_read": 3 }

The pipeline, one job per service

career pages fetch poolreal browsers tidy dedupe sort store API company graph · claims, not companies filter console · public site 1,000,000 pages a day, and nothing here reaches the web on its own 30+ services on this line

Schematic — shape only

MineLinkedIn and email outreach

Eight programs, two wires, nothing sent.

Started with

Outreach is eight small jobs: collect the lead, find who holds the role, write the message, produce the CV, sync the profile, sort postings, run the funnel, keep the ledger.

Runs as

Eight programs, one per job. Only two are joined. Every step that spends money is gated before each charge. The AI picks which facts to cite, never the sentence.

The call

Shut down the role lookup: every answer was a paid people-search call and the economics ended it, not the engineering. And an honest one: nothing has been sent. Either the gates work as built, or they’re set so tight nothing gets through. I don’t claim to know which.

For the engineer

The CV generator seeds the outreach engine with stored facts; the engine hands the message to an outside sender behind one seam. That seed and a replaced route to the role lookup are the only two wires, and neither is a call at runtime. The other five touch no sibling, reaching outward only.

bdr.local/gates
Spend gates and wiressender credential: not set
outreach enginewire: ← CV facts · → outside sender0 sent
CV generatormodel picks the facts, never the sentenceready
role lookuppaid call per answerclosed
lead collectorrented scrapers · gated before each charge€0 charged
posting triagereads the tabs open in your browserready
hiring funnelhalf its stations refuse to judge0 hired
profile syncwrites one live profile, re-reads afterready
outcome ledgerhands the campaign audience to the senderempty

Eight programs, two wires

CV generatorcites stored facts outreach engine0 sent role lookupclosed: economics sender seed replaced message — never delivered no wires to anything else inside lead collectorgated on spend posting triagesorts postings hiring funnelruns the funnel profile syncsyncs the profile outcome ledgerkeeps the ledger

Schematic — shape only

The thing you want to sell is probably the next one.

That’s the claim worth testing on a call: that the method transfers, not just that these worked. Twenty minutes with the engineer who did all six, and you leave knowing what I’d build first, what’s hardest, and what I’d talk you out of.

Book the twenty-minute call

Engineer this got forwarded to? Start at how it runs.

Mykola Vorobiov, photographed outdoors against forested mountains.
Mykola, Fredrikstad, Norway. The one who answers.