
Four language models, four Age of Empires 2 strategy scripts, from scratch, no peeking at each other. They failed the same way: hoard villagers, send hunting parties halfway across the map, never bank enough to advance an age. Gemini never left the Dark Age. Four vendors, one set of bad instincts.
The real find is what Emergent Garden built after: an overnight loop where a model mutates the best script, runs a tournament against the previous winner and the built-in extreme AI, keeps what survives. Genetic programming with an LLM as the mutation operator instead of a coin flip. ≈15 USD of tokens, and it beat extreme – by a hair, with a bot he then stomps by walking cavalry around the back.
Commenters point at the ceiling: the game ships Direct Unit Control – loops, pointers, near-human micro – and the models barely used it.

For a couple of months now every terminal I own wears the same bar: coralline, a Powerlevel10k-inspired statusline for Claude Code. All segments on, tokyo-night, two lines. Branch and dirty state, the active model plus its effort level, the context window filling up, the 5h and 7d rate-limit gauges with their reset countdowns, session cost, session length. I have stopped guessing how close I am to a wall.
Run the installer yourself, or hand the playbook to Claude Code and let it interview you. The readme’s own trust section says a Claude that stops to inspect that playbook first is behaving correctly. It is. Read the script, then run it – trust is good, verification is better.
Needs jq and a Nerd Font, or VL_ASCII=1 without the font. Renders locally: no network calls, no tokens.