A stack of pull requests, seventeen tall, towers over me. It’s eighty thousand lines of code, give or take. A total rework of an important component.
I rub my eyes, weary – it’s been a long day – and I type the magic word into the Slack channel: deploy.
The word isn’t magic in any real sense. It’s not picked up by any automated system, not watched for by any clever automation. No automated system may output it, either. It’s a manual mutex. It only works because every other human in the channel knows what it means, and agrees to play by the rules.
My mouse cursor hovers over the “Squash and merge stack” button. I’m feeling uneasy.
I read the pull requests, of course I did! Well… I skimmed them. It’s eighty thousand lines, and the very idea for this rewrite didn’t even exist this morning. I think I know what they do. “I think I know what I’m doing”, I mutter to myself and press the button.
As the seventeen pull requests merge, I rehearse the next steps in my head. I’m professional and diligent, so I prepared scripts to validate the new surface. I’ll wait for the production deploy command, run my scripts, and watch it all go green. I can already taste the green.
Tastes like metal for some reason.
My phone rings with an on-call page. Unrelated to the deployment,
minor, resolves itself in a few minutes. I groan and cd
back into the directory with my scripts. My focus is interrupted again,
this time by a Slack ping.
🚨 @pawel! PR #2137 contains a bug: ...
I got distracted, but the LLM-based release reviewer didn’t. It has already evaluated the deployment, and found three bugs. I don’t run my scripts. Not yet – clearly there’s some bugs in the code.
I could read the wall of text it output, digest it, and re-type it in
my own words. But that’s slower than the alternative. I click
Copy link to thread and paste it into the harness window
for the agent that coordinated building the stack of PRs that just
shipped. I follow it up, more out of habit than necessity, with “fix
plz 🙏”.
Then I read the wall of text myself. Only now, mind you; now that I expect a fix is already reasonably in progress. The bugs seem real, if minor: missed support for a parameter there, reverse sorting order than what I specified. Needs polish, not a revert.
By the time I’m done reading and digesting, the agent has already output a further stack of PRs, three tall this time. I skim it. Looks good. I type deploy into the channel again and press the requisite buttons, with the tired glee of a monkey that learned pressing the button labeled 🍌 causes bananas to happen.
The LLM reviewer found one more tiny bug – missing truncation of a
long string. I could fix it by hand, but I dutifully paste it into the
agent rectangle. For all I know, this might not be the only place with
broken truncation. By the time I’m done playing with
ripgrep and reviewing files, the agent has already checked
and produced a PR. It covers three spots: the one reported by the
reviewer, the one I found, and the one I missed.
I do the interpretive pantomime of shipping once more. The LLM review comes up clean. I run the verification scripts I prepared – well, that I had an agent prepare. They come up green. I mark the ticket as ready for review in production.
Minutes later, it is marked done. Another LLM reviewed it and deemed it unnecessary for a second human to take a look.
Did you think this was an AI doomer post?
Woe is me! Pity me, for I have become unnecessary! Nah,
sorry. This will not make for a good appendix to an anti-AI screed, I’m
afraid. The problem in that picture is not that there’s not enough of me
in it. It’s that there’s too much of me in it still.
I designed the rework in such a way that it could be deployed safely alongside current code. It was also designed in such a way that it is very easy to review automatically. Note that my own review plan didn’t involve meticulously checking it by hand. Instead, I had a script ready to go. I could’ve made that script – or the parameters for building one correctly – a deliverable. The LLM reviewer instead made one up on the spot, in a few seconds. It was correct enough to find bugs.
The problem, then, is that I had to type the magic word. I had to press the button. I had to copy and paste Slack links around. This, crucially, is not work. At least, it’s not work I should be doing.
Designing for automation
In designing the parameters for the rework, and the parameters for testing, I have arguably already done my job with this specific unit of work. Everything that came after could’ve been automated.
- The stack could have auto-deployed once CI went green,
- the auto-reviewer LLM could have been equipped to trigger fixes itself, rather than pulling my sleeve,
- these fixes could have auto-deployed too,
- the ticket could’ve self-marked as complete,
- and the final LLM pass could’ve happened exactly as it already has.
It’s even arguable I didn’t need to read or skim the code at all. We have code quality standards, lint, and test tools. I know that these were passing. If I was determined to read all of this myself and fully digest it, I would’ve delayed the release by another day at least. And as for the skim I did… it shipped with bugs anyway, so the outcome is the same whether I skim or not.
Of course, not everything is a clean, green-field-ish rewrite of a given surface. But we already came up with methods to make sure things are going well: canary cohorts, shadow deployments, A/B testing, feature flags. The tools were designed for a human to check things safely, but there’s no reason they don’t translate to an agent checking things safely.
They love yapping about “acceptance gates” anyway. I don’t know about you, but I’m feeling decidedly like a fence today – so my eyes should not be an acceptance gate input. The production logs should.
How do we avoid shipping bugs?
The unit of work isn’t a line of code, or a file, or even a pull request. It’s an idea. The path from an idea forming to it becoming a thing the customers can use needs to be as short as possible, and ideally involve as little human intervention as is necessary to make sure it’s working as specified.
That doesn’t mean “we must always ship things that are 100% correct”. Pro tip: very good attitude to have at NASA, not the best one if you want to move fast as a startup.
Sidekiq can deal with a Redis instance disappearing. An API client can deal with a 500 and a 429. We already build software with the assumption that things we interface with will break. I think software bugs are largely the same way: they are not the intended state, but they are the state that’s guaranteed to occur.
The question then changes. It’s no longer “how do we avoid bugs?”, because that has only one reasonable answer. We avoid bugs much the same way we avoid upstream provider downtimes – we don’t. We accept they are guaranteed to happen, and prepare the system to recover quickly.
The LLM reviewer didn’t need to wait for a human’s nod to solve a bug it’s already RCA’d. It could’ve just kicked off an agent to solve it.
Make it self-healing
When a customer reports a bug, is it better customer support to tell them:
Thank you! An engineer will take a look this week
Or:
Thank you! The system is already looking at fixing itself for you.
To be fair, I’m not satisfied with either of these answers. The one I truly want is:
Thank you! The system caught this five minutes before you started writing the e-mail. The fix has shipped two minutes ago. It’s already been checked against your reported regression. Go get ’em, tiger.
When there’s one correct decision, it’s not a decision
The correct decision regarding a report of “your deployment is buggy” is to fix it. The correct response to “a customer has complained” is to fix it. The correct response to “a set of changes has passed CI” is to ship it.
These aren’t hard and fast rules, but it’s a good enough heuristic to hold 99.9% of the time. The remaining percentage is absolutely when an automated system can and should escalate to a human – but a human deep in the loop is no longer the correct default. It should be a rare, carefully carved out exception.
It makes complete sense to me that what we should be building right now is not “the product” per se. We should be building the factory, the output of which is the product. Humans supply the important decisions, the so-called “taste”1. If my answer to an “A or B?” question is going to almost always be “A, obviously A, in what world is it not A, are you taking silly pills?”, then it shouldn’t have been a question in the first place.
Otherwise I’m not an important element of the process. I’m just a ceremonial human.