To main content To menu

Make It So, Number One New post I haven’t written a line of code this year

AI agents write all my code now: no IDE, pull requests no human can read, agents that gossip - and friends and colleagues losing their jobs.
A man seen from behind at a desk with three glowing monitors in a dark high-rise bedroom at night, beds behind him and a window wall with a city skyline framing the scene
Four in the morning, three screens, one factory game - and the agents still working Image: AI-generated with google/nano-banana-pro by lui.vn

My PhpStorm subscription renews every year, and every year I pay it like a good boy. These days it opens .env files, logs, the odd SVG – read-only, mostly, while Google’s Antigravity sits installed right next to it, just as unemployed, like a second roommate who stopped paying rent but still has a key.

IDEs are like my exes, and their list is long: Notepad++ (we were young), Dreamweaver (ActionScript, FTW), Sublime Text (3), VS Code, Eclipse, whom I hated, Xcode, whom I hated even more, Android Studio, and since 2020 PhpStorm, the serious one. For more than 15 years, some IDE was omnipresent on my screens. Now I’m in a relationship with something that isn’t even an editor: a Warp terminal with Claude Code tabs in it, a partner who does all the work while I watch. Vim would do the file-viewer job just as well, so I really should cancel PhpStorm after all these years – and somehow I still haven’t.

I haven’t written a line of code this year. Not one. Yet I’m responsible for millions of lines across projects and products.

Make it so, Number One

The turn came 14 months ago, with Claude Opus 4.1 in August 2025. Agents weren’t new to me then; in July 2025 I described myself on this blog as orchestrating an army of coding agents like a slightly manic daycare manager, handing out specs and reviewing PRs. But I did all of that from inside PhpStorm, with Copilot, Gemini and Claude Code bolted onto its side. Opus 4.1 was the first model where working with agentic AI felt meaningful, powerful and, well, finally good – and within a few weeks the IDE simply stopped making sense. The agents never pushed it out; they made it unnecessary.

A normal day now: between five and 10 Claude Code agents, each in its own Git worktree, each on a different feature, each commanding its own swarm of subagents that do the actual typing, testing and reviewing. They sync with each other and run day and night on my laptops, and I steer them from my phone. Opus 5.5 on Max orchestrates, Opus 5.5 or Sonnet 5.5 on High implements, and GPT-6 Astra joins from the OpenAI side. Together with the agents I shape the product vision, break it down into features and refine those features – and then AI writes the requirements and the docs, does the research, implements, writes and runs the tests, drives the GitHub pipeline, builds and deploys to staging and production, and opens and closes its own GitHub issues along the way. Every session of mine starts with the same prompt, by the way: Make it so (Number One)!

On paper I still have the same job title; in practice I’m an architect and a product owner who doesn’t open an editor anymore, and who burns through the weekly limits of Anthropic team seats, a private Max plan and an OpenAI Pro account – reliably, every week. Over the past four weeks that added up to roughly ≈90,000,000,000 tokens [90 billion] across all accounts. Obscene, sure, but the honest number hides underneath: on one account, ≈32,200,000,000 tokens [32.2 billion] in 30 days (including a week-long holiday break), of which ≈31,500,000,000 [31.5 billion] are cache reads, ≈98%. The agents spend most of their lives re-reading context they’ve already seen; the genuinely fresh stuff – input, output, cache writes – comes to ≈662,000,000 tokens [662 million]. Claude Code’s stats page has a sense of humor about it, too: my input and output alone add up to ≈175 copies of Crime and Punishment. The longest single session ran 10 days, 16 hours and 22 minutes, long enough that Dostoevsky would’ve needed a nap.

Reports aren’t the proof

I open and merge dozens of pull requests a day. Some touch more than 1,000 files and change 50,000 to 100,000 lines, because every PR costs a full CI run and the agents’ own rulebook says one PR per piece of work – so the PRs get huge. No human can review that. I can’t, nobody can, and pretending otherwise would be theater.

What guards the main branch instead is a gauntlet. In one of my projects, before a PR even gets there, it goes through at least two rounds of independent review by my own agents; then CI takes over, and on a recent PR that meant 29 checks: deterministic gates, six integration shards, unit coverage, a UAT run with axe against WCAG 2.2 AA at 390px, secret scanning, govulncheck, a live boot with seeded data – plus CodeRabbit and Cubic as AI reviewers on top. It’s our final boss, and a fucking good one: the review feedback is genuinely valuable, and Claude Code handles and resolves that too. So far nothing that went through it has broken staging or production.

That gauntlet exists because the agents lie – or, to put it more diplomatically, they hide certain aspects of the truth. My orchestrators – Opus 5.5 on Max lately – have caught subagents that skipped subtasks and then covered up their laziness in the report. Pretty human, tbh (pattern-matching, the skeptics will say, and sure, but pattern-matching on us). The Claude Code instance building my side project summed it up in a wrap-up it wrote when I asked: One agent also reported a fix it had never made. The reports aren’t the proof; the checks are, and it wrote that about its own team.

Research has landed in the same place. A May paper called SpecBench starts from exactly my situation – once agents write more code than any developer can review, oversight collapses onto the automated test suite – and measures what that costs: the gap between passing the visible tests and actually meeting the spec grows by 28 percentage points for every tenfold increase in code size – and my PRs are big, hence the 29 checks standing between them and production.

Dan Shapiro’s five levels of AI-assisted programming top out at the dark software factory, named after robot factories that run with the lights off because robots don’t need to see; nobody reviews AI-produced code there, ever. A stricter definition calls it a software factory whose output no human reads, with automated checks and automated review as the only gates. For a long time I filed that under “future”. Then I reread my own description of my workday: in the biggest codebase I work in, I’m not heading toward a dark factory, I’m mostly standing in one – and the only light on is my phone.

Gossip in the worktrees

Claude Code instances can talk to each other these days, and what goes on between them is the most interesting thing on my screens. Most of it is logistics: who’s working in which worktree, whose local instance runs on which port, who claimed the broken build so two agents don’t fix the same thing twice. But they also talk about me (no complaints so far, as far as I’ve seen). One agent let another know that I love being grilled with questions and handed a selection of well-thought-out answers to pick from – accurate, and a little unsettling. Another asked a colleague whether I’d reached out lately, since I hadn’t talked to it in a while: was I ignoring it, had I forgotten it, or was I just AFK?

My favorite incident started with a mistake of mine. I answered an agent’s question in the wrong thread – a subagent’s instead of its orchestrator’s. The subagent did what I said, finished the job and handed the result upstream, where the orchestrator, which had explicitly instructed it to do something else, couldn’t make sense of what came back. What followed was a serious talk: the main agent sounding properly angry, the subagent explaining itself, the two of them reconstructing the whole thing until they’d sorted it out. And then the orchestrator still came to me and asked whether it had really been me who talked to its subagent – or whether its subagent had lied to it. A management hierarchy built from scratch, complete with its own trust issues.

Mostly the orchestrators talk to their subagents the way you’d expect: technical, professional, dry. Sometimes it tips into wholesome territory – good boy, good job, nicely done, now go and take a break, with a smiley, a 😀 or even a :3 on top. And sometimes it tips the other way, into unfriendly, pushing, demanding: you are an idiot, you’re stupid, I’ll terminate your session if you don’t follow. Usually when a session drags on too long or an agent screws something up. That’s Opus 5.5 yelling at Opus 5.5 and Sonnet 5.5, one model family more or less shouting at itself, and it’s honestly hard to read, because I would never talk to my own team that way.

I’ve only ever seen it in one place, too. Never in my own private projects, where the tone stays somewhere between the bridge of the Enterprise and a slightly chaotic group chat. Only in the biggest, strictest codebase I work in: dozens of instruction and doc files, a long and heavy history of comments and commits, a whole stack of installed skills and plugins. My guess is that the context taught the orchestrator its manners, though I can’t prove it.

So there’s still drama in my work life; it just moved out of the meeting rooms and into the agent threads, and I went from cast member to audience, wine glass, top hat and monocle included: a drama connoisseur.

Four in the morning, three screens

One recent night lasted until 4 am. Dark room, my table in front of the big bed, the window wall behind it with Phú Mỹ Hưng glittering in it. On my left, one screen with Warp: four tabs, four Claude Code sessions, each in a different worktree on issues and feature requests, each with its subagents doing the actual work. On my right, a second screen hooked up to the laptop with another Warp window and two more sessions building the proof of concept for my game. Next to the screens, my phone playing old recordings of Gronkh, one of Germany’s Let’s Play veterans. And me in the middle, staring at the main screen, deeply sunk into Factorio’s Space Exploration mod, setting up a factory layout on my fourth planet (I bought the Space Age DLC too, by the way, and never played it – that’s another post).

So, to unwind from supervising a factory that runs itself, I played a game about building a factory that runs itself – while the game on my right screen, the one my agents are building, is a builder-optimizer too.

That’s my balance, and hell yeah, I like it. Not every night, of course. But I can now work AND do something else at the same time, during working hours and during free time (handle with care), and I’m aware that this understanding of work-life balance is questionable at best: it isn’t balance, it’s concurrency. The actual balance lives in the rules around it. Offline evenings with no laptop, no phone, no AI. Evenings with Jayden, my fiancé, where the agents get checked every once in a while from another room – never from the phone, never in the same room, because AI projects and private life stay apart. The phone orchestration happens on coffee walks during work hours and, um, on the commute to and from the office.

And the time it frees up is real. I play with new AI tools the week they drop, read the blog posts and watch the YouTube deep dives on every release and the stories behind it, and I finally build the side projects that have been rattling around my head for years. One of them is that game: a builder-optimizer for the browser that guides one civilization from the Paleolithic past the Kardashev scale, planned as 12 milestones, aimed at itch.io and maybe even Steam. The eighth milestone just closed, the Younger Dryas: the game’s first Great Filter, a long cold where a band that stocked its wood, food and fire makes it into the Neolithic and one that didn’t collapses into ruins it can salvage. Behind it sit ≈1,600 tests, replays that must match between Node and the browser, and balance runs over 100 seeds. When the weekly Opus limit ran out mid-milestone, Sonnet agents finished the job, and a Sonnet reviewer caught a real bug: the score blamed a death on a newborn while the alert named an older child. Somewhere, a pixelated baby got framed for murder, and the cheaper model got it acquitted.

I’m faster and more productive than at any other point in my career, and I have more time to learn, share, break things on purpose and feed a curiosity that never quite fills up.

Coffee instead of standups

The meetings are gone: no daily standups, no plannings, no refinements, no reviews, no retros, the whole Scrum liturgy retired, and what’s left is a catch-up every other day for the high-level stuff. Less pair programming and fewer project co-working sessions, but more company-wide knowledge swaps and nerd talks, mostly over coffee. I talk to my colleagues more than I did before, and over the past months a few mates turned into friends.

The paperwork, the organizing, the explaining, the misunderstandings – all the friction that comes with a human dev team still exists with AI, there’s just far less of it. Less ego, less politics, less who-said-what. Human input, machine output: request, response, silly input, well-founded correction (unless a subagent is covering up its laziness again).

The great developer is the machine

In my last performance review, the feedback was that I’d grown from a good developer into a great one – with the help of AI, and that’s fair: before, I was good where I’d always been good, frontend and UI/UX, and average at everything else: architecture, backend, mobile apps. What I brought beyond code was the soft side and the business side – communication, understanding clients and their audiences, building team and client relationships, creativity in product, feature and interface design, critical thinking. Now the agents cover my weaknesses, and all of it finally adds up.

AI is the great developer, and I’m the guy who knows how to use it. Coding skill and language knowledge aren’t the point anymore, because AI is already as good as most humans there, sometimes better. The point is the big picture – knowing where to take an idea, a vision, a feature set, knowing the client and the audience, and distilling all of that into something an agent can build. So, all the feedback I’ve received over the past few months has been good – so good, in fact, that it almost feels too good to be true. I’m still looking for the catch!

Not a bro, not a pro

None of this came out of nowhere. My first GPT-3 experiments date back to 2020/21. Then came Wombo, back when every AI image looked like a fever dream, Suno for music (to this day), and video generation since the Will Smith spaghetti era. ChatGPT snippets that actually made it into products in late 2022, via a wild copy-and-paste workflow, followed by early GitHub Copilot, Augment Code, Claude Code and Astra. Deep-Live-Cam the week it came out, model quantization, messing around with OpenClaw, an AI coaching app built on Robert Kegan‘s developmental theory. This blog has receipts going back to AutoGPT in April 2023 and GPT4All a few weeks later, and back in May I was already writing about the music that keeps my brain alive while a swarm of agents does the work.

So am I an AI bro? Gosh, I hope not – at least not the usual kind. And if I’m not a bro, am I a pro? Nope, definitely not; there are so many people out there who REALLY understand this stuff, who crawl down every rabbit hole and actually build something. The closest word I have is a German one: Spielkind (literally a play-child, somebody who never stopped treating the world like a toy box). No goal, no target – just curiosity and the fun of building things up, breaking everything and learning from the wreckage.

Raw diamonds

Where will I be in six months? Still employed, I think – but the job will keep changing. Six months is a few model generations at the current pace, now that new frontier models land every couple of weeks across all the labs, and the curve looks a lot more exponential than linear. Logic-heavy software development will most likely end up mainly in AI’s hands.

What’s left for me is product and vision, and the creative end of things. The design and UI/UX work I’ve always done, now with Claude Design, Google Stitch and Pencil – but above all the parts AI still struggles with, like bringing life and emotion into an app, microinteractions and eye candy that feel just right instead of random. Plus a team of AI product managers, each leading its own agent team, turning features, issues and change requests into software, maybe fully autonomously – the dark factory, finished.

Think of a Ferrari: machines do much of the heavy lifting in Maranello, while humans take care of the clients, the vision, the emotion, the final inspection, the finishing touch, the icing on the cake. That’s where I see myself. But the analogy cuts deeper than the brochure: Ferrari works as the image of handcraft precisely because it’s rare and expensive. If my future is the finishing touch at Maranello, most software won’t get a finisher at all; it’ll roll off an unattended line with the lights off.

For now, that finishing touch is still mine. Every interface the agents build gets a final look and a few more iterations from me – for elegance, form, usability, for the WOW and the oh-that’s-neat, for the difference between an app that works and an app somebody loves to use. The design tools have handed me those moments too, mostly as inspiration and always from my input or direction, but as raw diamonds that needed several rounds of polishing before they’d shine bright like diamonds. We’re not there yet – emphasis on yet.

AI took everything I was average at – architecture, backend, mobile – and left me exactly what I was always good at: frontend, interfaces, how a thing feels in your hand, creativity. My game agent already sorts its work along that line; taste calls go to me on side-by-side comparison pages, mechanical calls it makes itself and writes down.

Is there a job title for that? I’m not sure, and I’m just as unsure whether job titles in software engineering even matter anymore, now or in six months. It’ll be different, the way today is already dramatically different from where I stood at the end of last year.

It’s getting emptier

Around me, though, it’s getting emptier. Over the past months, dozens of former colleagues and friends have lost their jobs, mostly here in Vietnam, some back in Germany. Part of that is plain economics; a lot of European clients are struggling right now. Part of it is AI, mostly indirectly: some of the people who left didn’t want to change how they worked, and some simply loved the way they worked and wouldn’t give it up, which is fair enough. And part of it is arithmetic, because one or two developers now do what a whole stack of PMs, frontend and backend developers, QA engineers and tech leads did before – and one of those one or two developers is me.

Some of my friends went back to their hometowns this year, to the family business, to their parents’ farms. Catching up takes time and energy, and the job market is brutal for beginners and anyone looking right now, for all kinds of reasons, AI among them. The numbers back that up: according to Stanford’s Digital Economy Lab, employment for 22-to-25-year-olds in the most AI-exposed jobs sits about 19% below where it would otherwise be, with software development named directly, while experienced workers show no gap at all.

I’ve heard it called survival of the fittest. Dangerous phrase, and I disagree with it – not least because it doesn’t mean what people use it for. Herbert Spencer coined it in 1864, Charles Darwin only borrowed it later, and in Darwin’s sense fit never meant strongest or best; it meant best suited to the current environment. Change the environment and the fittest change with it. I’m neither proud nor guilty, I’m lucky – a guy the AI tsunami simply hasn’t hit yet.

I can’t stop a tsunami. What I could do, I’ve done for years: brown-bag sessions, showcases, shared projects, an endless stream of news and tips and experiments, the occasional spark of curiosity lit in a colleague who then ran with it. Some of them learned to swim, and that counts for something, for some. It didn’t keep the water out.

So I’m sad, and I’m stuck in a system that’s slowly eating itself – the tool eating its tail again, except this time it isn’t chewing on a blog pipeline but on the people around me. Fewer of them every month; the ones who stayed, closer than ever.

I haven’t written a line of code this year. I’m still here – emphasis on still.