CH 01 · Independent · 2026 – present · live product
A real Codenames game, and the AI is a player.
The premise
Codenames is a word game of clues and constrained guesses. Solo mode hands both sides of it to an AI.
It is deployed, it has players, and the AI is a participant rather than a feature bolted to the side of one.
AI Spymaster → AI GuesserIt plays both sides
In solo the AI is both roles. Only the clue passes between them.
AI spymaster · intent
BEIJING · WALL
AI guesser · calls
Discounted EGYPT — a weaker read alongside the other target.
Both rolesIt has to play legally
Each role can propose something illegal. Neither one gets to act on it.
Valid JSON is not a legal move — the validator decides.
Schema first, prompt second ↗(opens in new tab)Both rolesThe model underneath changes
A new model generation is not a dependency upgrade.
Looked like
A drop in the quality of the AI's play.
Actually was
An architectural gap. Contract checks and live telemetry surfaced it before it became a silent gameplay regression.
A migration ships when the contracts still hold under live traffic, not when the model is newer.
Model experiments as a stress test ↗(opens in new tab)Board · both rolesAnother language
The board just changed. The game did not.
Classic — one playable token per concept, in English and in Simplified Chinese.
Extended — authored per language, independent pools. No cross-language concept layer.
A language is supported when its own playable word set is written down and validated — not when its interface is translated.
This board is the extended zh-Hans pool. Its words are its own, not translations of the English board.
Whole productPeople actually play it
Not a demo. A deployed product emitting real sessions.
game_started
ai_clue_generated
ai_guess_generated
field_report_shown
reasoning_expanded
game_finished
Which sessions count is a question the instrumentation had to answer before any number was publishable.
Which sessions counted ↗(opens in new tab)CH 01 · Independent · 2026 – present · live product
A real Codenames game, and the AI is a player.
Codenames is a word game of clues and constrained guesses. Solo hands both sides to an AI.
They hit BEIJING, then WALL — used every call they were willing to make.
AI Spymaster → AI GuesserIt plays both sides
In solo the AI is both roles. Only the clue passes between them.
AI spymaster · intent
BEIJING · WALL
AI guesser · calls
Discounted EGYPT — a weaker read alongside the other target.
What reached the board
Both rolesIt has to play legally
Each role can propose something illegal. Neither one gets to act on it.
Valid JSON is not a legal move — the validator decides.
Schema first, prompt second ↗(opens in new tab)The player's side has not moved.
Both rolesThe model underneath changes
A new model generation is not a dependency upgrade.
Looked like
A drop in the quality of the AI's play.
Actually was
An architectural gap. Contract checks and live telemetry surfaced it before it became a silent gameplay regression.
A migration ships when the contracts still hold under live traffic, not when the model is newer.
Model experiments as a stress test ↗(opens in new tab)Board · both rolesAnother language
The board just changed. The game did not.
Classic — one playable token per concept, in English and in Simplified Chinese.
Extended — authored per language, independent pools. No cross-language concept layer.
A language is supported when its own playable word set is written down and validated — not when its interface is translated.
This board is the extended zh-Hans pool. Its words are its own, not translations of the English board.
Whole productPeople actually play it
Not a demo. A deployed product emitting real sessions.
game_started
ai_clue_generated
ai_guess_generated
field_report_shown
reasoning_expanded
game_finished
Which sessions count is a question the instrumentation had to answer before any number was publishable.
Which sessions counted ↗(opens in new tab)One product surface. Five different kinds of engineering behind it.
Qualified evidence
03 / live product
175+
monthly active players
Live-product telemetry · resume phrasing uses 175+ as durable floor, not the rolling monthly-active-players snapshot
#1
average position, branded search
Branded Google Search average position (last 28 days)
350
concepts in the classic word set
Classic set · word set × language → playable pool · one playable token per concept in English and in Simplified Chinese; the extended set is authored per language with independent pools, so equal pool sizes are a property of the classic set rather than a cross-language invariant
Where to go deeper
04 / optional depth
Experience it
Play Codenames AI
Solo runs both AI roles: an AI Spymaster proposes the clue, AI operatives call the board. Two Player keeps the spymaster human.
Solo vs AI · Solo vs JUDGE — safer, more cautious reasoning · Solo vs STRANGE — simulates outcomes before choosing · Two Player
codenames-ai.com ↗Read the reasoning
Four field reports
Inspect the build
Technical depth
Architecture and flow figures, the judge and validation path, the word-set model, and how this system connects to the others.
Private repository · no public URL