Case study · agentic software build

A spreadsheet macro became an airgapped reconciliation engine.
An agent fleet built the replacement - and put a brand-new model straight to work on it.

One report pair in, verified numbers out: 264 reconciliation rows against a 35,401-row general ledger, cell-for-cell parity with the verdicts the original macro produced. No human wrote the application code. This is how it was built, what broke, what it cost - and how a brand-new model got folded straight into live client work the week it shipped.

23.9h
active agent work, both frameworks
25
agent sessions, 5 models
175
tests green, parity gate included
264
hand-authored verdicts the engine must match

106 commits · 10 releases · 264 / 264 parity verdicts · 175 tests · a full usability pass on the deployed app · every figure metered from the logs, not estimated

01 · The challenge

Month-end close, held together by one macro and the one person who understood it

A senior ERP implementation consultant had handled his month-end reconciliation through a manual spreadsheet macro, refined over years of real close cycles - moving procurement and receivable records out of the systems that produced them and into the financial system. It did the job, and that was the problem: the logic lived in one file, in one head, and it could not tell a reviewer why a line had failed to match.

The brief was not "port the macro". It was: give the finance team the same verified numbers faster, with drill-through evidence the macro cannot produce, running entirely on their own laptop with no data leaving the machine - and make the engine general enough that the next report pair is a template, not a rewrite.

 The macro, as it wasThe engine, as it is
Who can run itThe one person who understands itAnyone with a browser and the two report files
Explaining a failureStatus only - the reviewer works out the reasonRoot cause plus the action the reviewer should take
Audit evidenceRebuilt by hand each month-endDrill-through to the exact source row
A second report pairRewrite the macroA template mapping - no code
Data leaves the machineNoNo — parsing, reconciling and export all run in the browser
Tests and version controlNone175 tests, two parity gates, full history
DistributionDepends on the person turning upA zip file that double-clicks

02 · The bet

No human writes the code. A fleet ships it, against a gate that cannot be argued with.

The build was delegated to a fleet of AI agents, with a human setting direction, making the product calls and running sign-off. It ran on two layers: a coding-agent framework that wrote the code, and a harness of orchestrator sessions around it that planned the work, reviewed it, deployed it and wrote this page. The workflow itself is the interesting part: plan with the most capable model available, stress-test the plan under interrogation, then execute with fast cheap models - and schedule the heavy work into the hours those models cost least.

STAGE 01 Plan & strategy qwen3.8-max The most capable model available plans, then refines under review. ~1.0h. STAGE 02 Scaffold & build glm-5.3-flash + muse-spark Parallel agents on one codebase. 20.2 active hours over three days. STAGE 03 Stress test deepseek-v4.1-flash Real client requirements, a model one day old. See chapter 07. 8 SEP 2026 12 SEP 2026

Stage 01

Plan & strategy

qwen3.8-max

Plan mode from an empty repo, then refined under interrogation until the plan locked. About an hour.

↓

Stage 02

Scaffold & build

glm-5.3-flash + muse-spark

Parallel agents on one codebase. 20.2 active hours over three days.

↓

Stage 03

Stress test

deepseek-v4.1-flash

Real client requirements, a model one day old. See chapter 07.

Three stages, six working days, five models - plus a harness of orchestrator sessions metered alongside. Every stage after planning ran on the same repository with the same gates.

The gate is the client's own numbers

Correctness was never left to opinion. A python golden port of the original macro, plus the 264 hand-authored per-row verdicts, form a parity gate. If the engine and the golden disagree, the engine is wrong. The tolerance for money is 0.005 and nobody is allowed to widen it.

Fresh agent per task, state on disk

Each task went to a newly spawned agent with a written brief. State lives in git, the ledger and the handoff journal, never in a model's memory - which is why the fleet changed models mid-project without losing a step.

03 · The roster

Who did what

Twelve builder agent sessions in the coding framework, plus thirteen orchestrator sessions in the harness that directed them - across five models and six working days, on one codebase nobody had to explain twice. Every session is accounted for in the meter - hours and cost, not anecdotes.

RoleSessionsModelResponsibilityMetered
Orchestrator13harness (multi-model)Plan review, briefs, review gates, client comms, deploys, verification, and this case study2.63h attributed
Planning & plan review1qwen3.8-maxPlan mode from an empty repo, then refined under interrogation until the plan locked~1h
Engine & app builders3glm-5.3-flashReconciliation engine, cockpit, views, ExcelJS export, close pack12.89h
Scoped subagents5glm-5.3-flashData-quality preflight, styled export, UI foundation, status-flow mapping3.78h parallel
Trial run, not adopted2muse-spark-1.3Benchmarked over one night of real build work; dropped the next morning in favour of the incumbent3.57h
Stress-test model1deepseek-v4.1-flashGR/IR close pack, built on the model's release day (see 07)release day
Review & verification-harnessParity gates, layout bleed suite, mojibake gate, hosted route checksindependent
Not every try earns a place. muse-spark-1.3 was benchmarked the same way - a night of real build work on this repository, when it was one of the newest models out. It did not make the cut: slow to respond and over-talkative in practice, so the build went back to the incumbent the next morning. Its hours stay in the ledger, because a case study that only reports the models that worked is a brochure. Two new models were tried here. One is now a fixture; the other is a footnote. That ratio is the honest one.

04 · The drama (all true, all in the git history)

Where it got interesting

Six moments worth telling, because a case study that only lists successes is marketing, not evidence. Each one changed how the fleet works permanently.

What brokeWhat happenedWhat changed for good
The deploy that took the preview down The release archive included the server's own configuration file. Extracting it overwrote the live configuration and took the hosted preview offline. The config is excluded from every archive, always - and the rule is written into the agents' own conventions file so no future agent can repeat it.
The file that would not die A page removed from the build kept being served, because extracting an archive over a directory replaces files but does not remove the ones that are gone. Removed paths are now deleted on the server explicitly as part of every deploy.
Caught by the client, not the tests One build carried mojibake - accented and symbol characters re-encoded into garbled multi-byte sequences. The client saw it before we did. A mojibake gate is now a release step: the built bundle is scanned for broken sequences, and the served page must still contain its intended glyphs.
The editor that ate the characters A PowerShell round-trip silently re-encoded the guide files, leaving 42 replacement characters in each one. Editing non-ASCII content through that path is banned outright; every agent in this repo is told to use the repo's own tools instead.
The gate that refused to bend When the engine disagreed with the workbook's hand-authored verdicts, the tempting move was to relax the tolerance until it passed. The repo's first rule forbids exactly that. The tolerance stayed at 0.005 and the engine was corrected - the only reason the client's reviewers trust the output.
The orchestrator crashed mid-run During the stress test the agent host process died with no exit path logged. Nothing was lost. Durability came from files, not from the process staying alive: the meter had already written its samples to disk, the session resumed a minute later, and the release still landed inside the same hour.

05 · What was built

The product

The engine was validated against real data by the consultant whose macro it replaces - over multiple review rounds, not a single sign-off - and it now reconciles a report pair it was never written for.

AudienceWhat they get
The reconciliation team A cockpit that answers "what is out, how old is it, and why" on one screen: pivots by account and period, exception registers with drill-through to the exact source rows, aging, smart-match proposals for unkeyed journals, and a period-over-period compare that separates genuinely new exceptions from ones that carried over.
The reviewer's audit file A styled workbook that mirrors the macro's own formatting expectations, plus the close pack: a reconciliation statement per line, an exception log ordered by root cause with owner and aging, and an accrual certification workpaper with control totals and sign-off slots.
Privacy and distribution No backend at all. Parsing, reconciliation and export run in the browser, so the ledger never leaves the laptop and the whole product is a zip file that double-clicks - with a gated hosted preview for demos that needs no source handover.
The next report pair The engine is template-first: the first pair is a template, not the architecture. A build-time wizard refuses semantically nonsense mappings that are structurally complete but wrong, so the next client's pair is a mapping exercise.

06 · For the tech nerds

Technical overview

Vite + React 18 + TypeScript single-page app, zero backend, zero framework style systems. The interesting parts are the parse-time column pruning, the four-pass engine, and the parity discipline that holds it all to the client's numbers.

LedgerMatch system architecture: a reconciliation analyst loads two report files into the offline React SPA, whose TypeScript engine - gated cell-for-cell against a Python golden port and 264 hand-authored verdicts - drives the ExcelJS export and IndexedDB persistence, with an optional HMAC-gated hosted preview

The whole system on one map: the analyst's laptop boundary holds everything that touches the ledger; the parity harness runs only at build time; the hosted preview serves the same bundle but never sees client data.

FILES report pair in CSV / XLSX ingest prune to template columns, keep serials dq preflight headers, blank refs, duplicates, fx mix engine four passes, control identity, exceptions analysis aging, composition, segment pivots export 5-sheet styled workbook PARITY GATE — every stage above is held to it Python golden port of the original macro + the workbook's 264 hand-authored verdicts. Cell-for-cell. Tolerance 0.005, never widened.

Input

Report pair in

CSV or XLSX, parsed in the browser. A 235-column GL export is pruned at parse time to the 40 columns the template maps.

↓

Checks

Ingest & data-quality preflight

Headers, blank references, duplicates, FX mix. Values stay strings; Excel serial dates are preserved exactly as the macro saw them.

↓

Engine

Four passes

Aggregate → net by segment → receivables → ledger exceptions. A control identity ties the totals.

↓

Output

Analysis & export

Aging, composition and segment pivots, then a five-sheet styled workbook.

Held to it

The parity gate

A python golden port of the original macro plus the workbook's 264 hand-authored verdicts. Cell-for-cell. Tolerance 0.005, never widened.

35,401 GL rows against 235 columns parse in about a second in a browser, because only the template-mapped columns are kept. Charts are hand-drawn SVG and the stylesheet is a single theme file.

4 passes
aggregate → net by segment → receivables → ledger exceptions
264
hand-authored verdicts acting as an independent oracle
3
runtime heavies: React, SheetJS/PapaParse, ExcelJS. Charts are hand-drawn SVG
0
backend processes. The bundle size is a product feature: it has to double-click offline

07 · The economics

How much work went in, and when we ran it

23.9hof agent work, measured
  • What this chapter is really about. Not what the build cost - it ran on subscription plans, so the outlay is not a clean number - but how much work went in and when it was run. Those are measurable, and they are what makes a fleet affordable.
  • Active hours, not wall clock. Consecutive message gaps under 15 minutes count as work; anything longer is idle, and sessions left open overnight are not counted. A session's span on disk can look five times larger than the work actually done in it, so the union of real working windows is what gets reported - never the raw elapsed time.
  • Two frameworks, one ledger. The build ran in a coding agent framework, and the orchestrating harness around it - reviews, deploys, client comms, this page - ran in a second. Both are metered. Harness hours are pro-rated to this project by how much of each session actually concerned it, so unrelated work is not claimed as build time.
  • Rates move by the hour. Providers now charge roughly double for the same model in their busy window. This build's heavy work was scheduled outside those hours, which is the single biggest lever on what a fleet run costs.

Active agent work, by day

Hours of genuine work across both frameworks - consecutive message gaps under 15 minutes. Idle overnight gaps removed.

0 4 8 12 3.93 9.94 7.05 0.33 1.72 0.86 8 SEP 9 SEP 10 SEP 11 SEP 12 SEP 14 SEP
8 SEP
3.93h
9 SEP
9.94h
10 SEP
7.05h
11 SEP
0.33h
12 SEP
1.72h
14 SEP
0.86h
ModelSessionsActive hoursShare of build
qwen3.8-max — planning1~1.0h
glm-5.3-flash312.89h
glm-5.3-flash (parallel pool)53.78h
muse-spark-1.323.57h
deepseek-v4.1-flash52.55h
harness sessions — orchestrator132.63h

Share-of-build bars are scaled to the largest model total. Harness hours are pro-rated to this project by measured session relevance. No cost column here on purpose: this run rode subscription plans, so per-model outlay is not a meaningful number.

The stress test

A new model landed mid-build. We put it to work in the hours it costs half.

On 12 September the consultant returned his answers on a new deliverable. The same day, a newly released model - deepseek-v4.1-flash, then about a day old - was handed the job as a live stress test: build the root-cause classifier, the four-sheet export and the aging buckets against the real requirements baseline, on the real repository, with the parity gates already green. It went on to carry the case-study work too - 2.55 hours across both frameworks, not the single session a builder-only meter would count.

Newer models are also where the money hides. Most providers now bill by time of day, and a model that is cheap when nobody else is working is a different proposition from one that is cheap on paper. So the fleet's heaviest work was deliberately scheduled into the discounted window - and every hour this model ran on this project landed there.

Peak hours

$0.30

per million input tokens - the listed daytime rate

Input$0.30
Cached input$0.006
Output$1.20

Off-peak - where this build ran

$0.15

per million input tokens - exactly half

Input$0.15
Cached input$0.003
Output$0.60
100%
of this model's project hours ran in the discounted window
5.7h
of work on this model, every hour of it off-peak - verified from session timestamps
2x
what the same work would have cost in peak hours
cheapest
per active hour of any model in the run - the reason it was worth testing
Stress-test meterValueEvidence
Session window58 minutes10:28 to 11:26, single session
Active build time0.97hmessage-gap metering
Model spendunder $0.50list-rate equivalent at the discounted window
Commits landed3classifier, close-pack export, release v0.1.15
Tests after174 greenup from 146 at the time; 175 after the UAT fixes
Total across both frameworks2.55h5 sessions: 0.97h builder + 1.58h harness
Fair attribution matters: this model did not build the whole application - it arrived a day before that deliverable did, and the model table above shows the split. What the stress test demonstrates is narrower and still useful: a just-released model, given a real requirements baseline and a repo with a hostile test suite, shipped a client deliverable inside an hour, for under fifty cents, without breaking a single parity test. That is the claim, and the meter backs it. The scheduling point is the one that compounds: the same discipline that made this build affordable is what lets a small team afford to keep trying new models.

08 · Validated

Not "the agents finished it" — "the numbers still match"

Speed is easy to demonstrate and worthless on its own. What makes an agent-built reconciliation engine worth putting in front of a finance team is that it can be held to a standard it cannot talk its way around: the client's own verdicts, reproduced cell for cell. Everything below is a gate, not a claim.

GateWhat it provesResult
Parity gate The engine reproduces the 264 reconciliation verdicts the original macro produced, row for row, from the client's own workbook. This is the gate the build could not pass by being fast or plausible. 264 / 264
Automated suite Engine arithmetic, matching passes, exports and view logic, run on every change. 175 / 175
Cockpit parity Headline figures in the shipped cockpit against the reference numbers, checked through the real UI rather than the unit tests. 3 / 3 figures
Domain review The consultant who built and has run the original macro reviews the output every round, on real close data — not a one-off sign-off at the end. Every round
Independence The engine now reconciles a report pair it was never written for. Passing on the original pair could mean it merely copied the macro's behaviour. Second pair
15
app views walked end-to-end with real data
4
routes verified, gated and public
5
issues found in UAT — all fixed before sign-off
0
console errors or failed requests in the full pass

The usability pass was run against the deployed app, not a local build, and it was run to find problems — which is why it did. Five issues surfaced; all five were fixed and redeployed before sign-off. The most instructive one was the least technical: a brand-new user opening the app was shown a green “pre-flight ready” badge before loading a single file, because “no errors” and “passed” had been treated as the same thing. Nothing in the automated suite caught it, because every test asserted the flag with data present. A human opening the app cold caught it in seconds.

That is the shape of the whole project. The fleet is fast, and speed is not the asset. The asset is a standard precise enough to fail against — then fixing what it catches, in the open, before anyone has to find out in production.

Finally, the client's own reconciliation consultant ran the app on his work and confirmed it. Sign-off came from the person who owns the numbers, not from the team that built the thing.

09 · What we'd tell you

If you're wondering "could this work for us"

Where fleets shine

Work with an oracle: something that defines "correct" so precisely a machine can be held to it. This project had 264 hand-authored verdicts and a golden reference. That is why the output is trustworthy - not because the agents were careful.

Where humans stay essential

Taste, scope calls, and knowing what the client actually means. Every requirement in this build came from a human conversation; the fleet executed them and refused to guess.

The discipline is the product

Review gates, isolated environments, an append-only journal, and release checks that catch mojibake and offline deploys. Remove them and the same fleet ships plausible-looking code nobody should trust with a ledger.

Provider-portable, and schedulable

Five models worked on this codebase without a rewrite, because state lives in files and tests, not in any vendor's context window. Models are swappable capacity - and schedulable capacity. Running the heavy work when rates are low is the difference between affording one experiment a quarter and running several a week. The gates are still the asset.

See it for yourself

The product page describes the engine. If you want the detail behind this page - which models, what they cost per hour, how the work gets scheduled, and how the gates are built - get in touch.