michael@multipliers.dev

Projects / evidence index

Every system, and what it is allowed to claim.

Two case studies carry the argument: a production AI product operated since 2026, and experiment measurement and attribution at Atlassian. Three more systems are where a requirement led further down — workflow, governance, and agent-native tooling. Each row states its contract and the evidence behind it.

Index key

Co-primary — written up at length, figures on this page

Supporting — one line, proof surface linked

Infrastructure — method, not a headline project

System / eraContractQualified evidence

CH 01

Codenames AI

Independent
2026 – present

Open →

A live game people play, where the AI’s move is legal before a player ever sees it.

A public AI-assisted Codenames, designed and operated end to end since 2026. A model proposes a clue and nothing in the reply guarantees it is legal on this board. Schema-first structured outputs plus board-aware domain validators separate valid JSON from a legal move; model migrations run as controlled experiments with evaluation built into product behavior rather than left as an offline afterthought. Supporting Chinese gameplay turned out to be a model problem rather than a UI one: a word set and a language resolve to a playable pool, so each language has its own playable word set over stable concept identity.

Structured outputsLanguage word setsModel migrationsTelemetry quality

Operated end to end — React/TypeScript, Node.js/Express, OpenAI, deterministic validation, analytics instrumentation, CI/CD, multi-environment deployment

350

concepts in the classic word set

Classic set · word set × language → playable pool · one playable token per concept in English and in Simplified Chinese; the extended set is authored per language with independent pools, so equal pool sizes are a property of the classic set rather than a cross-language invariant

175+

monthly active players

Live-product telemetry · resume phrasing uses 175+ as durable floor, not the rolling monthly-active-players snapshot

codenames-ai.com → · 3 field reports on the case study

CH 02

Experiment measurement

Atlassian
2020 – 2025

Open →

Nobody could defend the number until the system behind it was repaired.

Attribution windows and pipeline gaps silently change what an in-flight experiment appears to say. A Cross Flow funnel observability audit, an authored attribution formula adopted for Growth Experiment Impact Estimation, and embedded event-pipeline work restored attribution for experiments already running — preserving statistical validity without restarting them.

Attribution auditWindow reliabilityEvent pipelineAcquisition onboardingAdmin Hub

Role spine — SSE 2019–2020 · EM 2020–2024, experiment operations for teams running cross-product experimentation, AIM platform concurrent · SSE 2024–2025

>10%

prior-approach over-attribution exposed

Cross Flow funnel audit · formula subsequently adopted for Growth Experiment Impact Estimation

9%–41%

uplift variance across StatSig attribution windows

In-flight experiments · Statsig · the range is the finding

Five additional qualified figures on the case study (pipeline recovery, OKR attainment, Admin Hub activation). No field report is tagged to this work.

Supporting

Editorial workflow

Human-in-the-loop weekly field reports — retrieval, critique and verification before publish, with irreversible publish steps held by a person.

15 reports published via workflow · DEV series

Supporting

Renovate governance

Classifier, investigator and maintainer with distinct write boundaries, overridable stop causes, and merge authority held away from the proposer.

2 field reports · case study

Supporting

Agent-native systems

Team-harness plugin, cloud-hooks primitive and a four-layer hook stack — reusable workflows with explicit contracts and authority constraints.

Ecosystem walkthrough →