7 records
Findings
Measurements produced by the projects, with their method and sample.
- Games reaching an ending without the external diaryUnverifieddraftOnly twenty-one per cent of games reached an ending when the agent had no external memory to read back after its context compacted.
- Losses where an imminent rival victory was never checkedUnverifieddraftIn seven of twenty losses where a rival's victory condition was visible in advance, the agent did not check for it once in the twenty turns before losing.
- Plan follow-through, Claude Opus 4.6UnverifieddraftThe proportion of concrete next-moves Claude Opus 4.6 wrote down for itself that it actually carried out within ten turns.
- Plan follow-through, Gemini 3.1 ProUnverifieddraftThe proportion of concrete next-moves Gemini 3.1 Pro wrote down for itself that it actually carried out within ten turns.
- Plan follow-through, GPT-5.4UnverifieddraftThe proportion of concrete next-moves GPT-5.4 wrote down for itself that it actually carried out within ten turns.
- Share of agent actions spent checking the whole boardUnverifieddraftStepping back to assess the overall position accounts for only one to two per cent of what an agent does.
- Victory-condition checks performed against the sixteen instructedUnverifieddraftAgainst an instruction to check rival victory progress every twenty turns, models managed between four and ten checks in a 330-turn game rather than the roughly sixteen expected.