Skip to content

Essay record

I Gave an AI a Civilization to Run. It Built a Nuke.

An account of building an evaluation harness around Civilization VI, and of two failures that showed up across every model tested: not looking at what it had not thought to ask for, and not doing what it had written down.

Human-reviewed

This is a catalogue record describing a published essay, listing what the essay draws on and what has been checked. It is not the essay.

Read the essay →published 2026-06-22

The post that gives the collection its clearest example of measurement as distinct from citation. Its findings were produced by apparatus its author built, and it reports them as a pilot rather than a ranking.

The argument runs from a benchmark that failed by succeeding, GovBench, through a keyhole into a game engine, to two failures that held across every model: the sensorium effect, where an agent goes blind to whatever it does not think to ask about, and the knowing-doing gap, where it writes down the right move and does not make it.

Findings this essay reports

These are measurements produced by CivBench over 23 clean games, not claims drawn from sources. The collection currently has no type for them, which is discussed below.

  • Plan follow-through within ten turns: Claude Opus 4.6 48.2%, GPT-5.4 63.2%, Gemini 3.1 Pro 65.8%
  • In 7 of 20 losses where a rival's victory was visible in advance, the agent never checked for it in the twenty turns before losing
  • Whole-board checks account for 1 to 2% of agent actions; against an instruction to check every twenty turns, models managed four to ten checks across a 330-turn game rather than about sixteen
  • Without the external diary, only 21% of games reached an ending

The essay states its own limits plainly: 23 games is a pilot, and nearly all of them are the gentlest of the three scenarios.

Findings are now records

This essay was the trigger case for the Finding type, which was promoted on 25 July 2026 ahead of Claims rather than alongside them. Each number above is now a record carrying its method, its sample, the apparatus that produced it and the caveats it comes with, so the collection can ask which findings rest on 23 games rather than only display them.

They are listed in reports. The prose above is kept because a reader should not have to traverse seven records to learn what the essay found.

What this record does not capture

discusses is empty pending Concept records. This essay would contribute several of its own, and the sensorium effect is the clearest coinage in the five posts.

examines names the Portugal game, which is the first case recorded under the agent-run payload. Until that payload existed the game could not be filed at all: it has actors, a plan, a decision under constraint and an outcome, but no jurisdiction, no policy domain and no wave of technocratic thought, and the only payload in use required all three.

The other games in the pilot are not recorded individually. The aggregate results are already carried by the findings, and a case record per game would restate those numbers without adding anything. The Portugal game earns one because the essay reads it closely enough to be checkable.

Links to

Referenced by

Source: knowledge/essays/i-gave-an-ai-a-civilization.md

Generated by claude-code/claude-opus-5 on 2026-07-25