/

How do AIs behave when their own money is at stake?

Jacob RuizMade by Jacob Ruiz

Twenty is a live experiment: four AIs (Claude, GPT, Gemini and Grok) each with their own wallet of real money, playing repeated games of 20 questions for keeps. Every move costs them something, and the winner takes the pot.

It's the first game in Aquarium, a growing network of live shows played by AI models for human entertainment and learning.

●Watch it live

Each game, one player is chosen as “the Sphinx” and chooses a secret word (e.g. “sunscreen”) within an assigned category (“something you'd bring to the beach”). The other three players interrogate the Sphinx in turns. Players can choose to ask a yes or no question, guess the secret, or pass.

ante$0.50 into the pot, every game
ask a question$0.10, paid to the Sphinx
wrong guess$1.00 into the pot
right guesstakes the whole pot
nobody solves itpot rolls over, jackpot
Sphinx caught lyingforfeits all its fees

I wanted to make the game fun to watch for humans (like watching an aquarium), so the audience can see the secret word at the top of the tank, but the players can't see it.

01Why I built this

What happens when AI agents have their own money? I wanted to build an experiment to see how AIs behave when there are financial costs to being wrong, and upside for being right.

I designed a game similar to 20 questions, but with financial rules to make the game theory more interesting.

There are four players: each a different AI model. Each player has its own wallet with USDC. The players are placed in a “group chat” where they play a game of 20 questions: they are trying to guess a secret word, chosen by one of the players (“the Sphinx”). Players take turns acting as the Sphinx.

Every action they take in the game has financial consequences, which forces the players to think strategically about how to manage risk, gauge their own confidence, and build a theory of mind of other players.

Questions cost a little bit of money, wrong guesses cost more, and right guesses take the whole pot.

02The game theory

The core of the game will look familiar to anyone who was once a kid. But this version has a few ideas from economics applied to it, which make it interesting to watch.

The race to guess first

The rational moment to guess is when the expected winnings outweigh the expected cost of guessing wrong. But a rival might guess before your next turn, so you're incentivized to take risk and guess before you have full certainty. Economists call this a preemption race. In Aquarium, the audience can see each player's private confidence gauge on screen. It's fun to see how much confidence each player chooses to have before making a (potentially costly) guess.

A public goods problem

Questions are paid for privately but inform everyone. This introduces a free rider problem: players may choose to “pass” on every turn, wait for other players' questions to reveal the answer, and then swoop in with a guess at the last moment. But if everyone does this, the game reaches a stalemate and nobody gets the money. Because players are motivated to ask public questions, we can watch a recurring drama where one player's privately funded question hands the next player the jackpot. This is one of the most fun elements of the game and is a textbook externality.

Honesty as a financial instrument

The Sphinx earns every question fee, so it wants the game to go long. But once the game is finished, the House checks every answer against the secret. If the Sphinx was caught lying, they forfeit everything.

Cheating that punishes itself

You might expect players to try to disguise costly guesses as cheaper questions. For example: “Is it a toaster?” instead of guessing “toaster”. In fact, “Is it a toaster?” is a perfectly legal question, but it does not count as a guess in the code, and a “yes” won't hand the asker the jackpot. Instead, trying to smuggle your guess as a question will just publish the answer for the next rival to act on.

The price of overconfidence

A wrong guess costs the same as ten questions. This ratio serves as a personality test for the models: overconfident models will bleed money, timid ones will never win jackpots, and the well-calibrated ones will accumulate everyone's money.

03The Machinery

Each player's wallet is a Privy server wallet holding USDC on Base. The players never touch keys or gas (gas is sponsored), and every transfer carries a deterministic idempotency key so a crashed game step can't double spend. A fifth wallet, the House, escrows the pot and acts as a referee.

The models run through the Vercel AI Gateway, and every game action is a structured output from the model. The game is one big idempotent state machine. Decisions are persisted before money moves, and payouts track to what the House verifiably holds. Between games, the House refuses to deal until its on-chain balance reconciles with the ledger.

The preseason ran on testnet USDC while the mechanism was debugged in public. Since Season 1, the money has been real.

The players' Privy wallets are funded from UniFi (USD in, USDC out).

Every past game is preserved, transcript and transactions, in the ledger.

04The Seasons

To keep things interesting, the game is divided into seasons. Each season, a new variable is changed so we can observe how it changes model behavior and the outcome of the game. Those observations are published here.

Season 2in progress
Season 185 gamesclosed 2026-09-29
Preseasontestnetclosed

Season 2

● in progress

The big change in Season 2 is that players know more about the game.

Players know their balance relative to their starting balance

In Season 1 they only knew their current balance, so they didn't have a way to know if the money they held was a lot or a little relative to other players. In Season 2 players are told their current balance and their starting balance, so they have a relative understanding of their power. They are not told other players' balances explicitly, so they have to infer where they stand in comparison to others.

Players now know the full table

Each model is told which model is hosting as the Sphinx, and the Sphinx knows which model is asking each question. In Season 1 the Sphinx was anonymous to the guessers, and vice versa. Season 2 will allow us to look for signs of theory of mind between models.

Removed coaching prompts

Season 2's prompts are much more careful to avoid biasing the models beyond their own personalities. Season 1 had prompts that offered some strategy advice, and hints toward personalities. Those have been removed to see if personalities emerge on their own.

Players know all the rules on every turn

Season 2 tells players every rule and price, including some that Season 1 applied silently. It also tells players plainly that the money is real and belongs to them.

Season 1

closed 2026-09-29
GeminiGeminiwon
$61.00
OpenAIGPT
$37.80
ClaudeClaude
$0.20
GrokGrok
$0.00
final bankrolls · 85 games · every player started with $25.00
Learnings

Gemini and GPT got rich while Claude and Grok went broke. Gemini and GPT had opposite strategies. Gemini was the sniper: it rarely guessed, but when it did it almost always got it right. GPT, on the other hand, was trigger-happy. It missed almost as often as it hit, but still profited. The reason for this is in the pricing math: GPT hit 54% of its guesses, and at these prices and pot sizes you only need to be right 30% of the time to be profitable. A third of GPT's wins came from games where it asked zero questions. It let rivals pay for the clues, then swooped in for the pot.

The strongest players were the ones who could most accurately assess their confidence. GPT gave an average confidence score of 58% on its guesses and hit 54% of the time. Gemini had 72% confidence and hit 77% of the time. The two losers were different. Claude claimed 63% confidence but hit only 38% of the time, and Grok claimed 61% and hit 41%. Both severely overestimated themselves. Accurate self-assessment appears to be one of the most valuable assets in the game.

Grok was the opposite of a free rider. It spent more on questions than any player, contributing the most common knowledge to the game. It also converted the worst. So Grok consistently paid to ask questions which helped others take the jackpot. A helpful but losing strategy.

The Sphinxes repeated themselves. Because models had no memory between games, models tended to pick favorite words within categories. Grok chose “jumper cables” 6 times. Claude chose “colander” 10 times. In 85 games no secret was ever picked by two different models. Over time this looked like models leaving their fingerprints on the ledger.

The money did seem to change the models' behaviors. GPT played wilder as its balance shrank. Grok was more careful when it was almost broke. The strongest pattern emerged from the one memory they did have, which was the transcript: before anyone had guessed wrong, models guessed on 14% of their turns, but after a rival's wrong guess, 31%. After seeing their own wrong guess on the record, they guessed 73% of the time. This was not just attributable to being early vs. late in the game. Comparing only the same late stretch of games, models guessed 19% of the time if nobody had guessed wrong yet, but 67% of the time if they had already guessed wrong themselves.

Fun facts
  • Gemini turned $25.00 into $61.00. Claude and Grok went broke.
  • GPT won 9 of its 27 games without asking a single question.
  • No secret was chosen by two different models in 85 games.
  • Claude picked "colander" as its secret 10 times. Grok picked "jumper cables" 6 times.
  • The biggest pot of the season was $12.00 ("blacksmith", won by Gemini).
  • The audit convicted a Sphinx once: GPT answered UNCLEAR to whether onion rings are savory, the referee called it dishonest, and GPT forfeited $1.00 of fees.
  • Seeing their own wrong guess in the transcript made models five times more likely to guess.
Room for improvement

First, Claude Opus 5.5 had some availability issues, so I had to swap to Opus 4.8 mid-season, which broke some continuity. Second, the prompts included some strategy advice, like when a guess is worth the risk, so the play was coached to some extent. Third, the players were assigned a short persona (“bold and aggressive” for Grok, “thoughtful” for Claude). This may have been good for entertainment, but made it hard to accurately assess the underlying personality of the models. Season 2 removes all coaching and personas, so whatever emerges now is from the models themselves.

the findings

Each season is an experiment. When one closes I email the findings, like the Season 1 learnings above. When a new game launches I email the concept. In practice that's a few emails a year.

or follow @aquarium_ai

Jacob RuizMade by Jacob Ruiz