Dev.to AI 🤖 Ai 👁 0 📖 4 min read

DominoPoker: Test-Driving Multi-Agent AI Development (99% Written by AI)

Can you build a complete, production-ready, real-time multiplayer game with a trained AI engine and an authoritative Node.js server using almost exclusively AI? For the past weeks, I have been testing the limits of mod

DominoPoker: Test-Driving Multi-Agent AI Development (99% Written by AI)

Main Lobby
Can you build a complete, production-ready, real-time multiplayer game with a trained AI engine and an authoritative Node.js server using almost exclusively AI?

For the past weeks, I have been testing the limits of modern LLMs. The result is Domino Poker — a complete, fully playable Progressive Web App (PWA) written 99% by AI.

Multiplayer Lobby
bidding
web-view room
global leaderboard
settings
profile
personal statistics

In this post, I want to share the exact multi-agent architecture, verification loops, and prompt-pipelining that made this project possible without writing the actual code myself.

The AI Core Team: Three Models in a Loop

Instead of relying on a single AI model to write, check, and commit code (which often leads to hallucinations, bugs, and regressions), I designed a multi-agent system where three top-tier models collaborate, review, and double-check each other's work.

Meet the Agents:

  1. The Architect & Main Worker: Claude Code (acts as the primary developer writing files and managing the workspace).
  2. The Code-Checker & Critic: Codex 5.5 (an advisory partner reviewing changes through MCP hooks).
  3. The Independent Reviewer: Gemini 3.5 Flash (acts as the secondary validator to resolve disputes and confirm stability).

Here are the two main execution loops I established for every single code change:

Scenario A: The Happy Path (PASS)

When an issue is successfully implemented on the first try:

graph TD
    A[Task Started] --> B[Claude Code writes code]
    B --> C[Send Git-diff to Codex 5.5 via MCP]
    C --> D{Codex decision?}
    D -- PASS --> E[Send Git-diff to Gemini 3.5 Flash]
    E --> F{Gemini decision?}
    F -- PASS --> G[Claude Code: git commit & push to GitHub]
  1. Claude Code receives a prompt and writes/modifies the files.
  2. Once done, it automatically packages the git diff and sends it to Codex 5.5 via MCP for review.
  3. If Codex 5.5 returns a PASS rating (confirming the code matches rules, has no obvious bugs, and fits the architecture), the diff is forwarded to Gemini 3.5 Flash for a final sanity check.
  4. If Gemini 3.5 Flash also returns a PASS, the pipeline automatically runs git commit and git push to GitHub.

Scenario B: The Refinement Loop (BLOCK)

When a bug, edge case, or architectural violation is detected:

graph TD
    A[Task Started] --> B[Claude Code writes code]
    B --> C[Send Git-diff to Codex 5.5 via MCP]
    C --> D{Codex decision?}
    D -- BLOCK --> E[Claude Code evaluates Codex critique]
    E --> F[Send critique to Gemini 3.5 Flash for verification]
    F --> G{Gemini agrees block is valid?}
    G -- Yes --> H[Claude Code fixes code]
    H --> C
    G -- No / Resolve --> I[Claude Code adjusts code or overrides]
    I --> C
  1. Claude Code writes the code and sends the git diff to Codex 5.5.
  2. Codex 5.5 spots a structural flaw (e.g., mixing single-player logic with the multiplayer database, a security risk, or memory leak) and returns a BLOCK with detailed criticism.
  3. Claude Code doesn't blindly accept the block. It sends the critique over to Gemini 3.5 Flash to cross-examine the feedback.
  4. If Gemini 3.5 Flash agrees the block is valid, Claude Code goes back to the drawing board, fixes the code, and submits the new diff to Codex 5.5.
  5. This loop repeats autonomously until all systems return PASS.

Why this Multi-Agent Loop

By pitting different LLMs against each other, the system mitigates the "sunk cost fallacy" that single models often suffer from.

  • Self-correction: A model that writes code is often biased toward its own implementation. Having a dedicated read-only critic (Codex) prevents lazy solutions.
  • Dispute Resolution: Using Gemini 3.5 Flash as an independent tie-breaker keeps the system from getting stuck in infinite loops where one model suggests a fix and another rejects it.

What the AI Built: Domino Poker Features

To test this loop, I didn't want to build a simple Todo app. I chose Domino Poker — a complex game featuring:

  • Trained ISMCTS AI: An offline single-player engine based on Information Set Monte Carlo Tree Search that simulates thousands of game states. It runs fully offline inside a Web Worker so it never blocks the Next.js UI thread.
  • Authoritative Multiplayer Server: A real-time Node.js backend using WebSockets, where all game states and economies are server-validated (Zod schemas), making the game fully anti-cheat.
  • Determinstic Game Shuffling: Driven by a seeded random generator, letting the server fully reconstruct any match state from its event log for auditing.
  • Live Chat Translation: Integrating server-side Google Cloud Translation, allowing players to communicate in their own languages (localized in 21 languages).
  • Responsive PWA UX: A Next.js 16 + React 19 web app with beautiful animated background themes (purchasable with in-game gold coins).

Conclusion & Lessons Learned

  1. LLM collaboration is superior to single-prompting. Distributing roles (Developer, Critic, Auditor) drastically reduces the bug rate in complex environments.
  2. The "Advisory" pattern works. Giving one agent the sole power to modify files (Claude Code) while keeping others as read-only reviewers (Codex, Gemini) keeps the workspace organized and avoids conflicting writes.
  3. Code generation is ready for complex systems. If you set up proper boundaries, validation scripts, and multi-agent review chains, you can write 99% of a production codebase using AI.

Have you experimented with multi-agent validation loops? Let me know your thoughts in the comments!


📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.