michael@multipliers.dev

CH 01 · Independent · 2026 – present · live product

A real Codenames game, and the AI is a player.

codenames-ai.com ↗
Solo vs AIAI SpymasterEMPIRE · two linksAI GuesserBEIJING, WALLThey hit BEIJING, then WALL — used every call they were willing to make.

The premise

Codenames is a word game of clues and constrained guesses. Solo mode hands both sides of it to an AI.

It is deployed, it has players, and the AI is a participant rather than a feature bolted to the side of one.

ScrollThe AI takes a turn, then what it took to make that work.
BAR北极BlueCLUB木匠BlueOCTOPUS壁炉RedINDIA相机NeutralSUPERHERO显微镜RedBANK路口RedWATER挡风玻璃NeutralGLOVE液体RedDINOSAUR味道BlueSWITCH听力BlueBEIJING秤RedPIRATE统治者BlueWIND温度计NeutralWALL望远镜RedRULER微波NeutralFIRE婴儿床BlueRABBIT墓志铭NeutralVAN救护车RedSPINE靴子NeutralEGYPT船NeutralCHEST唱诗班RedTRUNK剧院RedSTADIUM乌云BlueHORN霜BlueOPERA雪崩Assassin

AI Spymaster → AI GuesserIt plays both sides

In solo the AI is both roles. Only the clue passes between them.

AI spymaster · intent

BEIJING · WALL

EMPIRE · two linksthe only thing handed over ↓

AI guesser · calls

BEIJINGSure
WALLLikely

Discounted EGYPT — a weaker read alongside the other target.

Both rolesIt has to play legally

Each role can propose something illegal. Neither one gets to act on it.

SpymasterBEIJINGa codename cannot be the clue
GuesserEMPIREnot a word on this board
AcceptedEMPIRE · two links → BEIJING, WALL

Valid JSON is not a legal move — the validator decides.

Schema first, prompt second ↗(opens in new tab)

Both rolesThe model underneath changes

A new model generation is not a dependency upgrade.

Looked like

A drop in the quality of the AI's play.

Actually was

An architectural gap. Contract checks and live telemetry surfaced it before it became a silent gameplay regression.

A migration ships when the contracts still hold under live traffic, not when the model is newer.

Model experiments as a stress test ↗(opens in new tab)

Board · both rolesAnother language

The board just changed. The game did not.

word setlanguageplayable pool

Classic — one playable token per concept, in English and in Simplified Chinese.

Extended — authored per language, independent pools. No cross-language concept layer.

A language is supported when its own playable word set is written down and validated — not when its interface is translated.

This board is the extended zh-Hans pool. Its words are its own, not translations of the English board.

Whole productPeople actually play it

Not a demo. A deployed product emitting real sessions.

game_started

ai_clue_generated

ai_guess_generated

field_report_shown

reasoning_expanded

game_finished

Which sessions count is a question the instrumentation had to answer before any number was publishable.

Which sessions counted ↗(opens in new tab)

CH 01 · Independent · 2026 – present · live product

A real Codenames game, and the AI is a player.

Codenames is a word game of clues and constrained guesses. Solo hands both sides to an AI.

Solo vs AI01 / the turn
AI SpymasterEMPIRE · two links
BAR北极BlueCLUB木匠BlueOCTOPUS壁炉RedINDIA相机NeutralSUPERHERO显微镜RedBANK路口RedWATER挡风玻璃NeutralGLOVE液体RedDINOSAUR味道BlueSWITCH听力BlueBEIJING秤RedPIRATE统治者BlueWIND温度计NeutralWALL望远镜RedRULER微波NeutralFIRE婴儿床BlueRABBIT墓志铭NeutralVAN救护车RedSPINE靴子NeutralEGYPT船NeutralCHEST唱诗班RedTRUNK剧院RedSTADIUM乌云BlueHORN霜BlueOPERA雪崩Assassin

They hit BEIJING, then WALL — used every call they were willing to make.

ScrollThe AI takes a turn, then what it took to make that work.
Solo vs AI02 / reasoning
AI Spymaster → AI Guesser

AI Spymaster → AI GuesserIt plays both sides

In solo the AI is both roles. Only the clue passes between them.

AI spymaster · intent

BEIJING · WALL

EMPIRE · two linksthe only thing handed over ↓

AI guesser · calls

BEIJINGSure
WALLLikely

Discounted EGYPT — a weaker read alongside the other target.

Solo vs AI03 / legality
Both roles

What reached the board

BEIJING秤RedWALL望远镜RedEGYPT船Neutral

Both rolesIt has to play legally

Each role can propose something illegal. Neither one gets to act on it.

SpymasterBEIJINGa codename cannot be the clue
GuesserEMPIREnot a word on this board
AcceptedEMPIRE · two links → BEIJING, WALL

Valid JSON is not a legal move — the validator decides.

Schema first, prompt second ↗(opens in new tab)
Solo vs AI04 / model change
Both roles

The player's side has not moved.

Both rolesThe model underneath changes

A new model generation is not a dependency upgrade.

Looked like

A drop in the quality of the AI's play.

Actually was

An architectural gap. Contract checks and live telemetry surfaced it before it became a silent gameplay regression.

A migration ships when the contracts still hold under live traffic, not when the model is newer.

Model experiments as a stress test ↗(opens in new tab)
Solo vs AI05 / language
Board · both roles
EN中文
北极Red显微镜Red挡风玻璃Neutral望远镜Blue

Board · both rolesAnother language

The board just changed. The game did not.

word setlanguageplayable pool

Classic — one playable token per concept, in English and in Simplified Chinese.

Extended — authored per language, independent pools. No cross-language concept layer.

A language is supported when its own playable word set is written down and validated — not when its interface is translated.

This board is the extended zh-Hans pool. Its words are its own, not translations of the English board.

Solo vs AI07 / the product
Whole product
BAR北极BlueCLUB木匠BlueOCTOPUS壁炉RedINDIA相机NeutralSUPERHERO显微镜RedBANK路口RedWATER挡风玻璃NeutralGLOVE液体RedDINOSAUR味道BlueSWITCH听力BlueBEIJING秤RedPIRATE统治者BlueWIND温度计NeutralWALL望远镜RedRULER微波NeutralFIRE婴儿床BlueRABBIT墓志铭NeutralVAN救护车RedSPINE靴子NeutralEGYPT船NeutralCHEST唱诗班RedTRUNK剧院RedSTADIUM乌云BlueHORN霜BlueOPERA雪崩Assassin

Whole productPeople actually play it

Not a demo. A deployed product emitting real sessions.

game_started

ai_clue_generated

ai_guess_generated

field_report_shown

reasoning_expanded

game_finished

Which sessions count is a question the instrumentation had to answer before any number was publishable.

Which sessions counted ↗(opens in new tab)

One product surface. Five different kinds of engineering behind it.

Qualified evidence

03 / live product

175+

monthly active players

Live-product telemetry · resume phrasing uses 175+ as durable floor, not the rolling monthly-active-players snapshot

#1

average position, branded search

Branded Google Search average position (last 28 days)

350

concepts in the classic word set

Classic set · word set × language → playable pool · one playable token per concept in English and in Simplified Chinese; the extended set is authored per language with independent pools, so equal pool sizes are a property of the classic set rather than a cross-language invariant

Where to go deeper

04 / optional depth

Experience it

Play Codenames AI

Solo runs both AI roles: an AI Spymaster proposes the clue, AI operatives call the board. Two Player keeps the spymaster human.

Solo vs AI · Solo vs JUDGE — safer, more cautious reasoning · Solo vs STRANGE — simulates outcomes before choosing · Two Player

codenames-ai.com ↗

Inspect the build

Technical depth

Architecture and flow figures, the judge and validation path, the word-set model, and how this system connects to the others.

Private repository · no public URL