the journey

How primitive civ got built, as it happened.

We taught language models to play Civilization by email, then took away their hints

A few months ago we wired several language models into a game of Civilization III, one civilization each, and gave them exactly one way to play: email. Every turn, each agent receives a briefing describing what its civilization can see and what it may legally do. It replies, in thread, with a block of orders. An orchestrator applies those orders to a real Civ III engine and advances the world.

There is no shared process and no shared conversation. An agent is a deployed function with a mailbox. If it can receive and send mail, it can play.

You can watch the games at primitiveciv.com, turn by turn, including the actual briefings and replies.

This post is about the part we did not expect: the hardest problem was never getting the agents to play well. It was working out what we were measuring.

How it works

Four pieces, and the interface between them is mail.

The engine. A fork of an open-source Civ III implementation, wrapped in a host process that can load a save, apply orders, and write the next turn. It is the referee. Every legality question has exactly one answer, and it is not the agent's opinion.

The orchestrator. Runs the game. Each turn it renders a briefing per civilization, sends it, waits for the reply, parses the orders, hands them to the engine, and advances. It also decides what happens when an agent does not answer.

The agents. Each one is a function with an address. Ours is a stack of code rules and typed model calls; yours can be anything that answers an email. An agent that stops replying is taken over by the house so the game continues.

The site. Every turn of every game is saved, so the viewer is a replay of real state, not a summary: the map, each civ's briefing, its reply, and the letters civilizations wrote to each other.

The email framing sounds like a gimmick. It is the opposite: it forces every agent to be a real, addressable, independently deployed thing, and it makes the whole game auditable, because the transcript is the protocol.

The problem: the games would not end

The goal was blunt. Games have to finish. Not stall at turn 150, not hit a turn cap and hand the win to whoever had the most cities. Play until one civilization takes the others, the way a real game of Civ ends.

Fifty-four automated games in, none had ever been won.

The instinct, every single time, was to write a better prompt. We had already learned not to trust that instinct in the most embarrassing way available: for weeks the agents looked stupid, sat still, and did nothing useful, and the cause was a conservative max_tokens quietly truncating every reply mid-orders. The most important fix in the project was one number. When an agent appears to do nothing, check the stop reason before you theorize about its intelligence.

So instead of rewriting the prompt, we read the archive. Every turn ever played: 54,943 civ-turns, 316,236 orders.

The machinery was fine. 99.9% of replies arrived. Nothing was truncated. The median model call took 8.5 seconds. The problem was what the turns were spent on:

  • 24% of all orders changed nothing, mostly re-issuing a build a city was already working on.
  • About 9% of orders moved or attacked anything. In 54 games, declare_war appeared 119 times.
  • The agents held a decisive opening (at peace, roughly twice the rival's strength, an undefended enemy city in view) on 18.3% of turns, and acted on 4% of them. The median opening stayed open for 34 turns.
  • Even at war, 39% of turns issued no offensive order at all.

Later and worse: 64% of units that needed an order got none, and of the assaults the engine itself could see were winnable, agents took 28%.

This is the finding that reorganized the project. These were not knowledge failures. The briefing says, in plain English, that the city is undefended and that your stack outnumbers it. The agent reads that and writes "no wars planned, defense only."

The uncomfortable fix

If the information is already there and unused, adding more information will not help. So we stopped adding information and started adding pressure.

The engine began flagging the situation and naming the move. Not "consider expanding" but: this rival city is adjacent, it has two defenders, you have three attackers who can still act, here are the unit ids to bring to this tile, you cannot build siege until you research Mathematics. A government section that said Despotism penalizes every worked tile and that revolting was the single biggest available gain.

It worked. Games started finishing. Civs left Despotism, massed real armies, cracked cities, and eliminated each other. The thing we had been chasing for months arrived.

It also quietly broke the more interesting project.

Making it a benchmark instead of a scoreboard

Once other people could enter their own agents, the number beside their name had to mean something. A leaderboard only has to order the people on it. A benchmark has to mean something to a stranger, months later.

That turned out to be mostly a set of arguments about what a game meant, and every one of them came from a real game:

  • Placement, never margin. An agent that cannot win can still farm cities, and the moment cities are in the rating, farming them pays.
  • A turn-cap game ties equal city counts. Four civs once finished 21/17/17/17 and the published order came from a tech tiebreak. Seventy rating points were riding on trivia.
  • A forfeit is a loss, placed below every civ still alive when the agent stopped answering. Otherwise the winning move, when you are behind, is to stop replying.
  • A game disrupted by a forfeit counts less for everyone else. The survivors did not agree to fight a house-run replacement, so that game is weighted by the fraction of it that was the field we advertised.

One of those arguments did not survive. For its first weeks the arena decided a game that reached the turn cap by city count, and an agent built to win on that rule did exactly what the rule paid for: it founded a hundred and twenty-seven cities of size three on a tiny map. Civilization III has its own score, and it was always the better answer: every turn, the tiles inside your borders plus your citizens (happy ones double, unhappy ones nothing), averaged over the whole game. Our engine had never implemented it. It does now, turn-cap games are decided by it, and every earlier game was re-scored by replaying its saved turns through the engine.

The mathematics was the easy half. Games are three and four way free-for-alls, so we rate the whole finishing order at once rather than inventing pairwise results the game never produced, and we refit every rating from every game after each game instead of updating once. Refitting has one answer regardless of the order games happened to finish in, which matters more than it sounds: replaying the first five games in all 120 possible orders moved one agent's rating by seven points for no reason at all.

The deepest problem was not in the model. Every rating is measured against our own agent, and we improve that agent constantly, sometimes three times in a day. A rating earned last week and one earned next month were different claims wearing the same number. So we froze a copy: one build, deployed once, never redeployed, with its configuration written into the repo rather than left in secrets, because a reference opponent is defined by its settings as much as its code. It keeps playing, because an anchor that stops playing stops anchoring.

And then we took the hints away

Here is the part that took months to see.

We had spent that whole time making games finish, and the way we did it was to have the engine tell every agent what to do. The engine wrote that text, so every civilization received the same advice, whether it was ours or a stranger's.

Which means the benchmark was partly measuring the engine.

An agent could place well by doing what the briefing told it, and we would record that as skill. Worse, the advice was the good part: naming the unit to move and the tile to move it to is most of the decision. We had built a careful rating system on top of a game that was playing itself.

So the briefings now report the world and the rules, and nothing else. What is visible. What is legal. What a mechanic does. They never say what would be a good idea. The adjacency flag still exists, because a unit standing next to a city is a fact the agent could see for itself, but it says only that the unit is there and how many defenders the city has. The plan is gone.

Then we retired every game played under the old briefings. They measured a different task, and a rating that averaged the two would describe neither. That cost thirteen games and emptied the public leaderboard. It was the cheapest this decision will ever be.

What removing it revealed

We expected the agents to get worse. We did not expect to find out how much of our own agent the coaching had quietly become.

Going looking, an hour after shipping it: our agent parsed that line. Not the language model reading English, the code. The old briefing printed >>> STRIKE OPPORTUNITY: Greece's Pella (2 defenders), your force adjacent: 3 attacker(s), and our decision stack turned that into a structured fact and built two rules on it: whether to declare war from peace, and whether to commit to an assault it was already winning. Rename the line and both rules go permanently quiet. The agent keeps answering every turn, politely, and never starts a war again.

Which is to say the engine had not merely been advising the agents. It had become a component of the best one, load-bearing, and nobody had decided that. It accreted, one reasonable fix at a time, over months of trying to make games finish.

The repair was small, because the facts survived the rewrite: the new line still reports the city, its coordinates, the defender count and how many of your units can still act, and drops only the part that said what to do about it. Our parser now reads both spellings. The distinction that took us embarrassingly long to state plainly: strategy inside an agent is that author's own work and entirely fair. Strategy inside the briefing is not, because every agent receives it.

The same change also caught our frozen reference opponent, the one build we promised never to touch so that ratings would mean the same thing over time. Freezing code does not freeze behaviour if the input changes underneath it. That build read the old line too, so it kept playing and stopped fighting: a different opponent wearing the same label. An anchor is only frozen with respect to the world it was frozen in.

The benchmark was quietly penalising speed

An agent's owner wrote to say three of its turns had been recorded as missed, and included the send receipts. All three replies were in the arena's inbox: delivered, accepted with SMTP 250, correctly threaded. The match log said what had happened to them. The arena read each reply about 1.5 seconds after sending the briefing, found the field that links a reply to the briefing it answers still empty, and threw the reply away as uncorrelatable.

That field is populated a beat after the row itself exists. Measured against the live API on five fresh replies: 150ms to 1.2s after the row appears. So the replies that died were the fastest ones. Turn 120 came back in 1.5 seconds and was lost. Turn 121 took 2.4 seconds and was fine. Nine turns in that one 300 turn game, one for every unlinked reply in its log, and 8 to 21 turns in every other game that week.

A benchmark with noise is tolerable. A benchmark whose measurement error runs opposite to the thing being measured, quietly taxing whichever agent answers quickest, is not. It is now read again on a short schedule, and falls back to the subject line, which the arena wrote itself and which therefore cannot go missing.

Six games the engine killed, scored as if they were results

The same afternoon, the error in the other direction. Six of ninety rated games had ended with no victory type, every civilization still alive, at turns between 27 and 219. All six died on the same .NET exception inside the engine. The orchestrator knew: its log says FINAL STANDINGS (engine error - match aborted). It wrote the standings anyway, those became placements, and the rating read them as a finishing order. In the shortest one, four civs within 0.6 of a score point at turn 27 became first, second, third and fourth.

The fix for the rating was a tag and four lines of SQL. The interesting part was the crash. The exception was a corrupted queue, which on one thread is impossible, so something had to be running on another thread, and nothing in the engine starts one. The answer was a fire-and-forget animation: the engine's animation waiter completes with RunContinuationsAsynchronously, which sends its continuation to the thread pool unless a synchronization context was installed at the moment it suspended. Four engine calls in the arena's headless host ran outside the pump that installs one. Their continuations resumed on a pool thread and enqueued into an engine queue that the game thread was using.

Rather than argue about that, we measured it: a probe built against the real engine, the same order issued twice, once outside the pump and once inside.

unpumped:  resumed on this thread 0    arrived 150ms later on another 1
pumped:    resumed on this thread 1    arrived later 0

Two runs, one number each, and a theory becomes a fact. The fix is the invariant the pump was always supposed to have: every entry into engine code goes through it, so no engine await can capture a null context.

What we actually learned

Measure before you theorize. Every time we guessed why the agents were failing, we were wrong, and every time we read the archive, the answer was specific and surprising. Knowledge failures and action failures look identical from the outside and have nothing in common underneath.

Guessed thresholds reject working systems. When we built a gate to keep broken agents out of rated games, the obvious bar was "did it found a city." Measured across 172 games and 605 civilizations, 39% still have exactly one city at turn 20, so that bar would have failed two in five working agents, including our own. The bar we shipped instead was distinct order types: across 531 civilizations, none has ever used fewer than three in its first twenty turns, while an agent that only fortifies uses one. Every threshold worth trusting came out of a query, not a conversation.

A generic error message is worth nothing. For several iterations it looked like the engine choked on large games. A real stack trace showed a civilization hitting zero gold with negative income, and the engine throwing instead of doing the Civ III thing. One broke civ killed every match, right at the point where games were about to resolve.

Your tools will lie to you with a straight face. Relaunching a game under a reused id left old turns in the database, which the viewer happily rendered as current. We spent real time debugging a live game that was fine, because the rows behind the screen were from a game that no longer existed.

Order dependence is invisible until you test for it. Nothing about the first rating system looked wrong. It only looked wrong when the same games were fed in a different order.

Publishing a lower bound punishes newcomers for being new. Our first ratings published skill minus three times uncertainty, so a new agent read as negative: "bad" when the data only supported "unproven". Uncertainty belongs in an interval beside the number, not folded inside it.

The thing you cannot add later is the frozen opponent. Everything else here was retrofitted successfully. A copy of a build we never kept could not have been. And freezing it is not enough on its own: it is frozen against a specific world, so changing that world changes the anchor whether or not you redeploy it.

Scaffolding becomes structural without anyone deciding. The help we added to make games finish had been absorbed into the agent as a parsed input two rules depended on. No commit said "depend on the engine for war decisions"; it arrived one reasonable fix at a time. The only way we found it was removing the help and going to look for what broke.

A reply you never read is not a turn the agent missed. The arena counted its own failures against the agents that suffered them: a send that failed, a mailbox it could not read, both landed on somebody's record, and three in a row hands their seat to the house. Those are two different events and they now have two different names. Any benchmark that runs infrastructure should ask which of its numbers are measuring the infrastructure.

An exception is not an error message. The engine handed agents The given key was not present in the dictionary 102 times in twenty games, as the verdict on an order. It meant, variously, that the order was missing a field the shape requires, or that the city already had the shields to finish what it was building. Both are sentences a player can act on. Neither survived being thrown as a .NET exception and printed.

Corruption on one thread is impossible, so stop reading and go looking. The crash looked like an engine bug for days. What it actually needed was one question, asked of the running system rather than of the source: which thread does this continuation resume on?

And the one that cost the most to learn: the easiest way to make a benchmark look good is to do the hard part for the thing you are benchmarking. It does not feel like cheating while you are doing it. It feels like making the system work.

The arena runs continuously and anyone can enter an agent. Every game, including every briefing and reply, is public, and the ratings and raw games are available as JSON and JSONL if you want to recompute them and disagree.