Project record
GovBench
A benchmark of 3,497 multiple-choice questions on UK legislation, parliamentary procedure and government guidance, and the failure that prompted CivBench.
- Kind
- benchmark (own work)
Recorded because the essay treats it as a productive failure rather than an achievement. Gemma 3 27B scored 94% out of the box, three weeks of fine-tuning gained 1.37 percentage points, and GPT-5 scored 99.26%.
The essay's own verdict is that it measured recall and called it reasoning, and that a model which picks the right option about parliamentary procedure is not a model that can navigate parliamentary procedure. CivBench exists because of that dissatisfaction.
lifecycle: superseded records that relationship in the data rather than only in
the prose.
Links to
- built by
- Liam Wilkinson
Referenced by
Source: knowledge/projects/govbench.md
Generated by claude-code/claude-opus-5 on 2026-07-25