About the bot
At the “Neural” level you play a neural network with 1.86 million weights. Nobody showed it how to play well. It only knew the rules and played 33 million hands against itself.
- 33M
- hands of self-play in training
- 62 min
- training on one graphics card (RTX 4090)
- 1–2 ms
- thinking time per move, in the browser
- 55.3%
- of hands won against “Strong”
Three bots with rules, one without
The Easy and Medium levels follow rules of thumb: lead cheap cards, keep pairs together, give up high cards only for tricks worth it. Strong adds a search on top. For each move it samples hundreds of possible distributions of the unseen cards, plays each one out and picks the move with the best average. Once no more cards are drawn it calculates exactly.
Neural does not search. The network looks at the position once and gives a probability for every legal move. Its judgement comes from training, not from rules a person wrote down.
What the network sees
A real example from self-play. Diamonds are trumps; the hand holds the king and queen of trumps and the trump nine, and the ace of trumps lies face up. The network gets no pictures, only numbers: for each of the 24 cards, whether it is in the own hand, already played, face up or still unknown.
Your hand
Face-up trump
Stock: 8 cards
Points 14 : 0, tricks 2 : 0
The card inputs: six rows of 24 cells
Each cell is one card. Filled means 1, empty means 0. Trump is always on the left.
| ♦︎ | ♣︎ | ♠︎ | ♥︎ | |||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| A | 10 | K | Q | J | 9 | A | 10 | K | Q | J | 9 | A | 10 | K | Q | J | 9 | A | 10 | K | Q | J | 9 | |
| Own hand | ||||||||||||||||||||||||
| Already played | ||||||||||||||||||||||||
| Shown by opponent | ||||||||||||||||||||||||
| Face-up trump | ||||||||||||||||||||||||
| Card led | ||||||||||||||||||||||||
| Still unknown | ||||||||||||||||||||||||
On top come 44 numbers for stock size, score, tricks, closing, the meld obligation and pairs in hand.
The network’s answer
Probability for each legal move. Value of the position: +2.9 game points.
One trick makes learning easier: the cards are rotated so that trump is always the first suit. Whether hearts or spades are trumps changes nothing about the game, so the network learns each situation once instead of four times.
Architecture
The network is deliberately simple: an input layer, three residual blocks and two heads. Each block fans the position out to 768 features, passes them through a GELU curve and compresses them back to 384. The shortcut around each block keeps training stable.
The policy head scores all 30 possible moves: 24 cards, 4 melds, exchanging the nine and closing. Illegal moves are removed before the choice. The value head estimates how many game points the hand will end with. That number is what the evaluation chart in the game shows.
In the browser the network runs in plain JavaScript, with no library and no server. The weights are stored as 8-bit numbers inside the game file (2.5 MB). The rounded weights choose the same move as the original in 99.5% of positions.
How much accuracy does rounding cost?
| Bits per weight | same move as the original |
|---|---|
| 8 | 99.5% |
| 6 | 97.6% |
| 5 | 95.0% |
| 4 | 89.8% |
36,000 positions from self-play.
Training by self-play
Training used PPO (Proximal Policy Optimization), a standard reinforcement learning method. In each of the 4,000 rounds the graphics card plays 8,192 hands at once. Half of them against the current network, the other half against up to 12 older versions, so it does not develop a weakness that only it fails to exploit.
The only feedback is the result: game points won or lost, 1 to 3 per hand. Which moves led there, the network has to work out itself. The learning rate falls from 0.0003 to 0.00001 over the run, so the network fine-tunes at the end instead of jumping around.
Win rate against the three rule-based bots during training
Measured every 200 rounds: 250 hands against Easy and Medium, 1,000 against Strong.
Values as a table
| Round | vs Easy | vs Medium | vs Strong |
|---|---|---|---|
| 200 | 75.0% | 59.4% | 50.6% |
| 400 | 76.5% | 60.1% | 52.3% |
| 600 | 77.1% | 62.3% | 51.5% |
| 800 | 73.2% | 60.7% | 53.2% |
| 1,000 | 76.6% | 62.4% | 54.4% |
| 1,200 | 77.3% | 65.1% | 53.6% |
| 1,400 | 73.7% | 61.9% | 53.1% |
| 1,600 | 78.1% | 62.1% | 53.9% |
| 1,800 | 78.7% | 63.3% | 53.6% |
| 2,000 | 76.2% | 66.0% | 52.7% |
| 2,200 | 78.0% | 64.0% | 54.5% |
| 2,400 | 77.0% | 60.9% | 54.3% |
| 2,600 | 76.0% | 61.5% | 55.3% |
| 2,800 | 77.6% | 66.3% | 54.7% |
| 3,000 | 74.8% | 62.7% | 55.9% |
| 3,200 | 75.9% | 63.0% | 56.3% |
| 3,400 | 76.2% | 64.5% | 55.6% |
| 3,600 | 78.0% | 64.5% | 57.1% |
| 3,800 | 74.8% | 64.0% | 56.8% |
| 4,000 | 77.1% | 65.2% | 56.6% |
What the network picks up: actions per hand in self-play
Average over both seats, measured every 10 rounds.
How strong is it?
Everyone against everyone, 1,000 hands per pair. Each deal is played twice with seats swapped, so luck of the deal cancels out. Read a row: how often this bot wins against the opponent in the column.
| row vs column, wins in % | Random | Easy | Medium | Strong | Net at 400 | Net at 1,600 | Net final (Neural) |
|---|---|---|---|---|---|---|---|
| Random | 19 | 8 | 7 | 5 | 5 | 5 | |
| Easy | 81 | 30 | 21 | 24 | 24 | 22 | |
| Medium | 92 | 70 | 36 | 39 | 38 | 36 | |
| Strong | 92 | 78 | 64 | 49 | 46 | 43 | |
| Net at 400 | 94 | 75 | 61 | 50 | 47 | 46 | |
| Net at 1,600 | 95 | 76 | 62 | 53 | 52 | 47 | |
| Net final (Neural) | 95 | 77 | 64 | 57 | 54 | 53 |
Does it play differently from the search bot?
Both win about equally often against the weak levels, but not in the same way. Measured over 10,000 hands per bot:
What did not help
After the first good network came many attempts to make it stronger: longer training, more inputs, a search on top, giving it the match score. None got measurably beyond the roughly 55% against Strong. The Monte Carlo search even made it weaker, because it pretends to know the opponent’s cards.
Two measurements explain why. A network trained for 16 million hands only to beat v3 reaches 49.8%: it finds no weakness. A network allowed to see the opponent’s hand and the stock, on the other hand, wins 77.8%. That gap is the price of hidden cards, not of weak play. Without peeking at other cards there is little left to gain.
Win rate per attempt, with standard error
Dashed: 50%, even. The bar around each dot shows the measurement uncertainty.
Against the search bot “Strong”
Against net v3 itself
The bot as a coach
The network helps you too. The Hint button shows its best move, how sure it is and what the second choice would be. When a rule of thumb agrees with the network, its reasoning is shown as well.
Under Training aids in the settings you can turn on the evaluation chart. Like in chess, it shows after each of your moves how the network rates the position, and grades your move: best, good, inaccurate or mistake, depending on how likely the network would have played it itself. At the end of the hand you see where it would have played differently.