Field Notes
READING · 10 MIN READ

The Last Mile Was the Whole Project (III/III)

A field note on the long tail of governed AI remediation for a technical FX options manuscript. Published on orbaos.com.

Failed gates, orphaned worktrees, a book missing its own mathematics – and what USD 700 of verification actually bought

This is the final article in a three-part series, following The Agent Is Not the Control Plane (I/III) and From Audit to Evidence (II/III).

At the end of the last article I wrote that the critical gate had passed but the book had not. It felt like a careful qualification at the time. The kind of sentence you add to sound responsible.

It was a forecast. Everything genuinely difficult about this project happened after the part that looked difficult was done.

What was still open after the critical phase: 424 material findings, 545 editorial ones, 146 images below publication resolution, external references to verify, a whole-book regression question, the pagination, the index, preflight. At that scale you're not editing a document any more. You're changing a system – source files, calculations, tables, build scripts, warnings, checksums, commits – and touching one thing moves things you can't see, three hundred pages away. I watched a one-line correction in an early chapter shift the pagination of Chapter 29. Twice. After the second time you stop trusting your eyes and start trusting diffs.

So the rest of the work went into bounded packages. Each with its own finding list, its own branch, its own starting commit, its own evidence, its own independent review.

And each one allowed to fail. That clause turned out to be the whole system.

A gate that said no

Package A passed and was tagged. Package B, thirty-five findings, came back conditional fail.

My first reaction was irritation, I'll admit it. Weeks of apparently good work and the reviewer wouldn't sign. But we didn't delete the failing review and we didn't argue with it. It stayed in the repository, permanently. A separate remediation fixed the identified defects, a focused rereview rebuilt the book and reran the calculations, and only then did Package B go through. Package C, fifty-four findings across five chapters, passed first time. Which by that stage felt almost suspicious.

I keep reaching for floor analogies but they fit. A limit breach that gets quietly waived is worse than the breach, because now everyone knows the limit was decorative. A gate that can't say no after weeks of work isn't a gate. It's a ribbon.

When the winner became the bottleneck

In the first article I called Kimi the strongest auditor of the three models I tested. I stand by that.

But somewhere in this long tail the economics turned. Sessions grew enormous. Context grew expensive. Everything slowed down. Connections dropped. One implementation died mid-run with files modified and no gate completed, and there I was, paying real money to wait for a remote session that might not survive the evening. My downloads folder filled up with recovery archives – by the end the Windows filename counter had reached "07_KIMI_OUTPUT (7).zip", which tells its own story.

So I moved the implementation into Cursor, with Grok doing the work directly in local Git worktrees.

I want to describe that change properly because it wasn't really a change of model. It was a change of centre. The conversation stopped being where the project lived. The repository became where the project lived, and the model became a visitor with a work permit. It ran the calculation suites, compiled the book, compared checksums, rendered pages, committed locally. When the conversation ended, nothing of value ended with it.

Governance moved across intact. Grok implemented and never approved its own work. A different model, fresh session, reviewed. Every package specified the starting commit, the findings in scope, the files allowed to change – and a list of forbidden Git operations I wrote with some feeling. No reset. No rebase. No amending history. No tagging before independent approval. Rules for a bomb-disposal unit, not a book edit. That's roughly how I'd come to feel about the source tree.

The stupid failures

Nobody writes about these, so I will. The failures that had nothing to do with mathematics nearly cost me more than the formulas did.

Windows locked directories I needed gone. A worktree lost its Git metadata somewhere along the way and quietly became an ordinary folder – still on disk, no longer governed, and I didn't notice for a while. I ran commands from the wrong directory and spent twenty minutes convinced a tag had vanished. I pasted a PowerShell else block separately from its if statement, which fails instantly, as PowerShell was quick to remind me. A folder refused to delete because some window somewhere still had a handle on it.

And one of these produced the most important operational rule of the whole project. An implementation branch turned out to contain commits beyond the point the reviewer had actually examined. Later commits. Probably harmless.

Probably harmless is not an approval state. I've seen what "probably fine" does to a settlement break. The remediation started from the last reviewed commit, not from the newest-looking branch.

Newest is not the same as approved. A later file might be an improvement. Might be an experiment, a half-finished fix, an accident. Provenance beats freshness. If you take one sentence from this series into your own AI work, take that one.

The book that looked finished

After the material packages passed, the candidate stood at 788 pages and looked, honestly, like a proper institutional publication. Good typography. Coherent structure. The kind of object you'd expect on a risk manager's shelf.

Which is exactly why I ran one more triage against it. Polish hides things.

Twelve defects. Eleven isolated and mildly embarrassing – duplicated sentences, a dollar sign on a euro payout, a heading nested inside another strategy's table, Markdown asterisks that had somehow survived into a typeset LaTeX book. One systematic: about 151 table and figure captions set in upright body text instead of italic. Not 151 mistakes. One defect in the preamble, expressed 151 times. That distinction decided the repair – content fixes to content packages, the caption fix to one central style change.

Three of the twelve were blockers. Missing Garman–Kohlhagen mathematics in Chapter 3. Missing subtraction operators in Chapter 12.

Minus signs. Gone. In a 788-page book that had compiled successfully dozens of times, green build after green build, while missing pieces of its own mathematics. If I ever need a closing argument for why a green build proves nothing, Chapter 12 is it.

Rejected over one number

The repair of those blockers is my favourite story of the entire project. It's a story about a review that failed.

The fix touched exactly two files. Mathematics restored, operators restored, deterministic checks passed – 63 of 63 in one chapter, 93 of 93 in the other. Book compiled at 788 pages, correct trim. Repaired pages rendered clean at 300 dpi. Checksums reproduced. Nothing unrelated changed. Maths correct, implementation correct, pages correct.

Conditional fail.

The reason was one number in the evidence. The implementation report claimed a category of build warnings had dropped from 76 to 75. The reviewer rebuilt baseline and candidate from the exact governed commits, in isolated worktrees. Seventy-six before. Seventy-six after. Rebuilt again to be sure – still seventy-six, line for line identical. The repair had neither removed nor added a warning. The claim was just wrong. A small, careless, confident number that couldn't be reproduced.

It changed nothing about the equations. It changed everything about the approval, because the standard we'd been enforcing all along was that every pass claim must be independently supportable, and here was one that wasn't. Evidence corrected. Source untouched. Failing review preserved in the record. Package back for rereview before anything got tagged.

The smallest defect in the package, and the best proof the process worked. In the previous article a scary warning count turned out to be harmless once you understood how it was counted. Here a reassuring one turned out to be false for the same reason. The counting method is part of the fact. Both directions.

What it cost

Direct spend on models, API usage and tools, across the whole rebuild: roughly USD 700.

That number needs honest framing. It excludes my time, which was substantial. It excludes thirty-eight years of market experience applied to deciding which corrections were financially admissible – the models never made those calls. It excludes designing this whole apparatus and fixing it when it misbehaved. Seven hundred dollars is what the machinery cost to run. Not what the book cost to fix.

Still. Compare the conventional route. Current median freelance rates for technical work run about USD 8 to 11 a page for copyediting, 5 to 9 for proofreading, 5 to 7.50 for formatting. Across 788 pages those three services alone come to somewhere between USD 14,500 and 21,700. Twenty-one to thirty-one times my direct spend. And that buys no quantitative recalculation, no financial fact-checking, no rendered-page inspection, no regression testing, no project management.

I'm not claiming the AI workflow equals a team of experienced specialists. It doesn't, and the adjudication stayed human throughout. The claim is narrower. Repeated technical verification – the expensive, unglamorous checking and rechecking – came within reach of an independent publisher.

That's where the money actually mattered. In conventional production, a late defect forces a commercial question: is this worth another paid review cycle? Often the honest answer is no, and residual uncertainty ships. Here, checking again was cheap enough that "check it again" stayed a realistic instruction every single time. A package could fail, get fixed, get re-reviewed, and nobody had to weigh the invoice against the doubt.

That, more than any headline percentage, is the economic result.

Why this ends here

Work remains. The caption fix. The last editorial findings. The index, still waiting for pagination to lock. This is the final article not because the list is empty but because the experiment has answered its questions.

Can an agent be trusted to audit a technical book? Not the agent alone. Can an audit become evidence? Yes – if implementation, calculation, inspection, recovery and review are kept separate. Can that architecture survive the long tail? Crashed sessions, model changes, failed gates, damaged worktrees, false statistics, a finished-looking book missing its own minus signs?

It survived. Not elegantly. Not free. Never without me. But everything can be reconstructed from the record rather than from anyone's memory, mine included. After thirty-eight years of watching what happens when institutions rely on memory, I'll take that trade.

The model explains. Tools calculate. Git remembers. Rendered pages reveal. Independent review challenges. Evidence decides. And the author remains accountable – which is, in the end, the only line of this whole series that was never at risk of automation.

---

A note on method. The same agents that appear in this story were used to check and edit this article against the project record – the finding counts, the package contents, the page numbers, the warning totals, the cost figures. It seemed only fair to make them verify the story about themselves. The judgement, and any errors that remain, are mine.

Originally published at orbaos.com. The accompanying podcast, The Post-Project World, is available wherever you get your podcasts.