Sechsundsechzig

About the bot

At the “Neural” level you play a neural network with 1.86 million weights. Nobody showed it how to play well. It only knew the rules and played 33 million hands against itself.

33M
hands of self-play in training
62 min
training on one graphics card (RTX 4090)
1–2 ms
thinking time per move, in the browser
55.3%
of hands won against “Strong”

Three bots with rules, one without

The Easy and Medium levels follow rules of thumb: lead cheap cards, keep pairs together, give up high cards only for tricks worth it. Strong adds a search on top. For each move it samples hundreds of possible distributions of the unseen cards, plays each one out and picks the move with the best average. Once no more cards are drawn it calculates exactly.

Neural does not search. The network looks at the position once and gives a probability for every legal move. Its judgement comes from training, not from rules a person wrote down.

What the network sees

A real example from self-play. Diamonds are trumps; the hand holds the king and queen of trumps and the trump nine, and the ace of trumps lies face up. The network gets no pictures, only numbers: for each of the 24 cards, whether it is in the own hand, already played, face up or still unknown.

Your hand

Face-up trump

Stock: 8 cards

Points 14 : 0, tricks 2 : 0

The card inputs: six rows of 24 cells

Each cell is one card. Filled means 1, empty means 0. Trump is always on the left.

♦︎♣︎♠︎♥︎
A10KQJ9A10KQJ9A10KQJ9A10KQJ9
Own hand
Already played
Shown by opponent
Face-up trump
Card led
Still unknown

On top come 44 numbers for stock size, score, tricks, closing, the meld obligation and pairs in hand.

The network’s answer

Probability for each legal move. Value of the position: +2.9 game points.

meld 40 (diamonds)48%
exchange the trump nine45%
play jack of diamonds6%
play nine of diamonds< 1%
play queen of diamonds< 1%
play king of diamonds< 1%
play ten of clubs< 1%
play king of spades< 1%
close the stock< 1%
The network hesitates between melding 40 and exchanging the nine first. Both are strong: exchanging brings the ace of trumps into the hand, and the meld is still possible on the next lead. It rates the other moves as almost worthless, such as giving away king or queen without melding.

One trick makes learning easier: the cards are rotated so that trump is always the first suit. Whether hearts or spades are trumps changes nothing about the game, so the network learns each situation once instead of four times.

Architecture

The network is deliberately simple: an input layer, three residual blocks and two heads. Each block fans the position out to 768 features, passes them through a GELU curve and compresses them back to 384. The shortcut around each block keeps training stable.

The policy head scores all 30 possible moves: 24 cards, 4 melds, exchanging the nine and closing. Illegal moves are removed before the choice. The value head estimates how many game points the hand will end with. That number is what the evaluation chart in the game shows.

In the browser the network runs in plain JavaScript, with no library and no server. The weights are stored as 8-bit numbers inside the game file (2.5 MB). The rounded weights choose the same move as the original in 99.5% of positions.

How much accuracy does rounding cost?
Bits per weightsame move as the original
899.5%
697.6%
595.0%
489.8%

36,000 positions from self-play.

What the bot sees188 numbers, trump always as the first suitInput layer 188 → 384Residual block× 3591,744 weights × 3LayerNormLinear 384 → 768GELULinear 768 → 384skipLayerNormPolicy head 384 → 30probability for each movelegal moves onlyValue head 384 → 1expected game points, −3 to +31,860,511 weights in total, stored as int8 in the browser

Training by self-play

Training used PPO (Proximal Policy Optimization), a standard reinforcement learning method. In each of the 4,000 rounds the graphics card plays 8,192 hands at once. Half of them against the current network, the other half against up to 12 older versions, so it does not develop a weakness that only it fails to exploit.

The only feedback is the result: game points won or lost, 1 to 3 per hand. Which moves led there, the network has to work out itself. The learning rate falls from 0.0003 to 0.00001 over the run, so the network fine-tunes at the end instead of jumping around.

Win rate against the three rule-based bots during training

Measured every 200 rounds: 250 hands against Easy and Medium, 1,000 against Strong.

40%50%60%70%80%even01,0002,0003,000training round77% vs Easy65% vs Medium57% vs Strong40%50%60%70%80%even02,000training round77% 65% 57%
After only 200 rounds the network clearly beats Easy and Medium. Against Strong it needs about 1,000 rounds to pull ahead. Single points vary by one or two percentage points. The final measurement over 4,000 hands: 55.3% ± 0.8.
Values as a table
Roundvs Easyvs Mediumvs Strong
20075.0%59.4%50.6%
40076.5%60.1%52.3%
60077.1%62.3%51.5%
80073.2%60.7%53.2%
1,00076.6%62.4%54.4%
1,20077.3%65.1%53.6%
1,40073.7%61.9%53.1%
1,60078.1%62.1%53.9%
1,80078.7%63.3%53.6%
2,00076.2%66.0%52.7%
2,20078.0%64.0%54.5%
2,40077.0%60.9%54.3%
2,60076.0%61.5%55.3%
2,80077.6%66.3%54.7%
3,00074.8%62.7%55.9%
3,20075.9%63.0%56.3%
3,40076.2%64.5%55.6%
3,60078.0%64.5%57.1%
3,80074.8%64.0%56.8%
4,00077.1%65.2%56.6%

What the network picks up: actions per hand in self-play

Average over both seats, measured every 10 rounds.

0.000.400.801.2001,0002,0003,000training round0.75 melds0.45 closes0.41 nine exchanges0.000.400.801.2002,000training round0.75 0.45 0.41
At first the network plays at random: it closes in 61% of hands and loses 43% of all hands that way. After a few rounds it almost stops closing and learns to meld first. Closing returns only afterwards, now on purpose: in the end almost every second hand is closed, and only 3% of hands are lost that way.

How strong is it?

Everyone against everyone, 1,000 hands per pair. Each deal is played twice with seats swapped, so luck of the deal cancels out. Read a row: how often this bot wins against the opponent in the column.

row vs column, wins in %RandomEasyMediumStrongNet at 400Net at 1,600Net final (Neural)
Random1987555
Easy813021242422
Medium927036393836
Strong927864494643
Net at 400947561504746
Net at 1,600957662535247
Net final (Neural)957764575453
The three networks come from the same training run. Already after 400 rounds the network is level with Strong; the final network beats every opponent in the table. It loses only 5 in 100 hands to random play, because even a random player sometimes gets the better cards.

Does it play differently from the search bot?

Both win about equally often against the weak levels, but not in the same way. Measured over 10,000 hands per bot:

Takes the trick when it canas second player, when a card would win
85%
76%
Leads an ace or tenshare of leads
30%
36%
Leads a trumpshare of leads
24%
23%
Trumps a non-trump leadas second player
24%
22%
Closes the stockshare of hands
37%
35%
Wins after closingshare of its own closes
92%
91%
The most visible difference: the network leaves a trick it could take more often (76% against 85%). In exchange it leads aces and tens more often. It saves trumps for later and prefers to collect the points actively. Both close about equally often and win afterwards about equally often.

What did not help

After the first good network came many attempts to make it stronger: longer training, more inputs, a search on top, giving it the match score. None got measurably beyond the roughly 55% against Strong. The Monte Carlo search even made it weaker, because it pretends to know the opponent’s cards.

Two measurements explain why. A network trained for 16 million hands only to beat v3 reaches 49.8%: it finds no weakness. A network allowed to see the opponent’s hand and the stock, on the other hand, wins 77.8%. That gap is the price of hidden cards, not of weak play. Without peeking at other cards there is little left to gain.

Win rate per attempt, with standard error

Dashed: 50%, even. The bar around each dot shows the measurement uncertainty.

Against the search bot “Strong”

Net v3, in the gamePPO with decaying learning rate
55.3%
Net v1first run, constant learning rate
54.8%
Net v4, six times longerbest checkpoint, re-measured on fresh deals
54.4%
Net v5, more inputssuit voids, played cards, last trick
56.2%
Net v6, knows the match scoreplays for the match, not for game points
55.6%
Net + exact endgame solvercalculates once no more cards are drawn
55.8%
Net + Monte Carlo search64 sampled deals per move
51.6%

Against net v3 itself

Dedicated exploitertrained on 16M hands against v3 only
49.8%
Net that sees all cardsknows the opponent’s hand and the stock
77.8%

The bot as a coach

The network helps you too. The Hint button shows its best move, how sure it is and what the second choice would be. When a rule of thumb agrees with the network, its reasoning is shown as well.

Under Training aids in the settings you can turn on the evaluation chart. Like in chess, it shows after each of your moves how the network rates the position, and grades your move: best, good, inaccurate or mistake, depending on how likely the network would have played it itself. At the end of the hand you see where it would have played differently.

Play against the network