How the Stage 11 engine behind Backgammon Sage Pro was built to handle the positions every bot plays worst — backgames, containment games and the snake — and what the benchmarks say about it.
Backgammon engines have always been weakest in the slow, structural games: backgames, where a player holds two or more anchors deep in the opponent's board and waits for a shot; containment games, where a player with a lost race tries to keep one or two hit checkers trapped; and the snake, a far-side prime holding a single straggler while the rest of the opponent's army crumbles. These positions are rare in ordinary play and unusually long when they do arise, so they are barely represented in the self-play data an engine learns from, and an engine that has scarcely seen them evaluates them by extrapolating from positions that merely look similar. The Open Sage Stage 9 engine was trained on backgame positions and scores better than eXtreme Gammon at matched evaluation levels, but its error rate in these games is still several times its error rate elsewhere. Stage 11 attacks the problem directly. We first built benchmarks made of backgames, containment games and snakes, each scored as a Performance Rating against rollout-grade references; then added six neural networks to the engine, selected by a small set of structural rules; then generated rollout targets for every one of them; and finally scored the result against every benchmark we have. On the containment benchmark Stage 11 cuts the 3-ply PR from 6.89 to 3.58 and the blunder count from 156 to 53; across the ten classic backgame benchmarks its 1-ply PR falls from 4.75 to 2.98; and its play in unlimited games and in Paskogammon — a variant played from a custom starting position that leads into far more backgames — is unchanged or better.
Neural-network backgammon engines learn from millions of self-played games, and they learn what those games contain. Backgames, containment games and primes holding a single straggler are rare in ordinary play and unusually long when they do occur, so a general engine sees few of them and values them by extrapolating from positions that look similar but are not. The result is a well-known weakness: an engine that makes an error every twenty decisions in a normal game can make one every three in a deep backgame, and it tends to misjudge the same things a human intermediate does — when to hold an anchor and when to run, how much timing a position has, and when a prime should be released.
The Open Sage Stage 9 engine, which powers Backgammon Sage Pro, already includes two networks trained specifically on backgame positions, and our bot performance study shows it evaluating at least as well as eXtreme Gammon at every matched level. But a strength relative to other bots is not the same as strength in absolute terms. Measured against rollouts, Stage 9's error rate in backgames is two to three times its error rate in ordinary games, and in containment games and snakes it is worse still: on the snake benchmark below its 1-ply Performance Rating is over 60, against a unlimited-game figure of 2.6.
Stage 11 is designed to remove those weaknesses without touching anything else. The work went in four steps. First we built a family of benchmarks that score backgame and containment play in different kinds of position — ten classic backgame types, a containment family, a massive-backgame family and a snake family — each a set of real decisions scored against rollout-grade references, so that Stage 9, Stage 11 and XG can all be measured on the same footing. Second we designed the Stage 11 model: Stage 9's seventeen networks carried unchanged, plus six new networks for the backgame families, and a routing algorithm that decides from the structure of a position which network evaluates it. Third we generated training data for every new network — a mixture of positions harvested from self-play, synthetic positions and existing rollout corpora, all rolled out to produce targets. Fourth we trained the networks and scored the finished engine against every backgame benchmark and against our rollout-PR benchmarks for unlimited games, five-point matches and Paskogammon (a variant played from a custom starting position chosen to produce a lot of backgames), to make sure the gains in backgames came at no cost elsewhere.
The short version of the results: Stage 11 is markedly stronger everywhere the new networks fire and identical to Stage 9 everywhere they do not. Its 3-ply PR on the containment benchmark is barely half of Stage 9's, its blunders there fall by two thirds, its pooled 1-ply PR on the ten classic backgames falls by 37%, and it is the first Open Sage engine to make real progress on the snake. Unlimited-game play is unchanged and Paskogammon play improves. The rest of this page gives the details.
Every benchmark here follows the same recipe as our unlimited-game rollout-PR benchmark: a set of real decisions, each with a rollout-grade reference evaluation, against which an engine is scored as a Performance Rating — its average equity error per decision, multiplied by 500, the convention XG uses. A PR of 2 means the engine gives away 0.004 of a point per decision on average; a PR of 60 means 0.12 per decision, which is the territory of a beginner.
What differs is where the decisions come from. Each backgame benchmark starts from a small set of seed positions that define its family — the reference positions shown in the carousels below. The engine then plays unlimited games out of those seeds against itself, cubeful with Jacoby and beaver, at 3-ply, and every decision that arises while the position still belongs to the family is recorded: a checker play counts when there are at least two legal moves and a meaningful equity spread between them, and a cube position counts when the doubler's decision is not trivially obvious or when a double is actually offered. Seeds are used in a shuffled cycle, which side moves first is a coin flip, and everything is a pure function of one random seed, so any benchmark regenerates identically.
Every recorded decision is then rolled out to produce its reference: 5,184 paths per position played to completion, with 3-ply checker and cube decisions for the first three half-moves and 2-ply thereafter, variance reduction on, distributed across a fleet of cloud workers. A checker decision's reference carries the rolled-out equity of every candidate the reference player's filter kept; a cube decision's carries the no-double, double/take and double/pass equities. One subtlety matters when comparing engines: a candidate outside the reference player's filter set is only ever valued at filter precision, which over-charges an engine whose picks differ from the reference player's. We therefore ran candidate-completion passes — rolling out, under the same convention, every move that Stage 9 or Stage 11 actually chose at 1-, 2- or 3-ply — so that every score below is against a rollout-graded candidate.
A backgame is named by the two points its owner holds in the opponent's home board: the 2‑1 backgame holds the opponent's 2- and 1-points, the 5‑4 holds the 5- and 4-points, and so on. Ten types cover the useful combinations. Each folder's seeds are hand-curated positions from XG reference files, and its benchmark holds about 1,000 decisions — recorded only while the two named anchors are still held and the holder is behind in the race — for a pooled total of 10,021 decisions. Every folder's seed positions are benchmark positions too, as cube decisions. Select a type to scroll through its seed positions; each card shows the position's rolled-out cube analysis alongside the board.
A containment game is defined by the escaper, not the container: one side has borne off a number of checkers and has one to three checkers that were hit and must run the whole board home, while the other side, with a lost race, arranges whatever it has left — anchors, blots, a broken prime — to keep hitting them, playing to save the gammon or occasionally to win. Nothing is required of the container's structure. The benchmark's 201 seeds come from positions the engine had already met (decisions from the unlimited-game and Paskogammon benchmarks with their cube, and rows of the backgame training piles), filtered by that rule, and its 3,197 positions — the seeds' own cube decisions among them — are recorded while the escaper still has trapped checkers. A second reference for a 300-position sample was rolled out with Stage 11 rather than Stage 9 playing the trials, to check that the choice of trial player does not decide the comparison.
A massive backgame holds three or more anchors in the opponent's home board, or two anchors with seven or more checkers back: the deep, timing-driven positions where the backgame player has committed most of the army. The 200 seeds are drawn from the same sources as the containment seeds, filtered by that rule and kept disjoint from the containment and snake families, and the benchmark holds 2,199 positions.
The snake is a far-side prime holding a single straggler: a run of four or more consecutive points, each made with two or more checkers, entirely on the opponent's half of the board, trapping a checker on the bar or in the holder's home board while the opponent's other ten or more checkers are already crunched home. It is a priming game rather than a backgame, and one of the hardest shapes in backgammon to play: the holder must roll the prime home without releasing the straggler, and the timing of the release decides the game. Because the position is rare in play, the 74 seeds include synthetic variants of the classic shape (prime length and placement, straggler location, crunch split) alongside the few found in existing data; the benchmark holds 1,073 positions.
Three benchmarks measure whole games rather than a family of positions, and exist to prove that nothing was lost: the unlimited-game rollout PR (500 games, 17,535 decisions), the five-point match rollout PR (130 matches, 18,292 decisions, with the match score threaded through every cube decision) and the Paskogammon rollout PR (2,556 decisions, played from Paskogammon's scattered opening position, which produces far more backgames and containment games than standard backgammon). All three build their references adaptively — 3-ply for clear decisions, truncated rollouts for closer ones, full 1,296-path rollouts for the closest — as described in the bot performance study.
XG has been scored on these benchmarks too. It exposes no programmatic interface, so every benchmark decision is written out as a one-decision game and replayed through its Batch Analysis feature, one pass per evaluation level, exactly as the unlimited-game and match benchmarks were. Those results are on the bot performance page, which compares the engine against XG; the tables here compare Stage 11 against the Stage 9 it replaces.
Stage 9 is a committee of nineteen neural networks: one for pure races, sixteen selected by the pair of game plans the two sides are pursuing (racing, attacking, priming, anchoring), and two backgame networks — one for the player holding a backgame and one for the opponent — used when the plan pair is anchoring against racing, the anchoring side trails in the race and holds two or more anchors in the opponent's home board. Every network is a single hidden layer of 400 units over 244 inputs (100 units over 196 inputs for the race network), producing five outputs: the probabilities of winning, winning a gammon, winning a backgammon, losing a gammon and losing a backgammon.
Stage 11 keeps the seventeen standard networks byte for byte and replaces the two backgame networks with six, giving a committee of twenty-three:
| # | Network | Region it evaluates |
|---|---|---|
| 0–16 | Stage 9's race and plan-pair networks | Everything else, exactly as in Stage 9 |
| 17 | Deep backgame | The 2‑1, 3‑1 and 3‑2 backgames: both anchors on the opponent's 1-, 2- or 3-points |
| 18 | Middle backgame | The 4‑1, 4‑2, 5‑1 and 5‑2 backgames: one anchor on the 1- or 2-point, one higher |
| 19 | Double anchor | The 4‑3, 5‑3 and 5‑4 backgames: two anchors, none deeper than the 3-point |
| 20 | Early containment | A backgame holder that has hit a shot while the racer still has at most two checkers off, in positions the plan-pair gate would otherwise hand to a standard network |
| 21 | Containment | Any containment game, whatever the container holds |
| 22 | Snake | A far-side prime holding a straggler against a crunched board |
Two design choices matter more than the count. The backgame networks are chosen by the category of the backgame — the same network whichever side holds it — rather than by which side holds it, so each learns one structure from both perspectives. And the network for a checker play is chosen from the position before the move, so a single network values every candidate, including the ones that leave its region (the move that gives up an anchor, or releases the prime); that is what makes hold-or-run decisions consistent, and it is also why the training data for every network has to include the positions those release moves lead to.
The routing algorithm is a short ordered list of structural tests on the position, the first that matches winning:
Because the tests are structural, their footprint can be measured on any set of positions, and that footprint is what bounds the risk: on the unlimited-game benchmark the containment rule fires on 1.3% of positions, the early-containment rule on 0.6% and the backgame rules on 2.4%, while the snake rule fires on none at all; on the Paskogammon benchmark the figures are 2.6%, 20% and 25%. The 95% of ordinary positions Stage 11 evaluates identically to Stage 9 are, by construction, unchanged.
Every new network was trained the way Stage 9's backgame networks were: supervised learning on the GPU against rolled-out targets — the five probabilities a full rollout assigns to a position — warm-started from an existing network, with the checkpoint that scores best on a held-out tenth of the data kept. What differs per network is where its positions came from, and each answer taught us something.
Stage 9 and Stage 10 had already accumulated rolled-out backgame positions — standard games and Paskogammon games, about 640,000 rows in all. Splitting them by category gave 245,000 deep, 167,000 middle and 94,000 double-anchor training rows, with held-out benchmark rows deduplicated against them. Trained on those alone, the trio was better than Stage 9 inside its region and worse at its edge: a network that has only ever seen positions with both anchors held has no idea what the position is worth once an anchor is given up, and since it values every candidate of a backgame decision, it guessed — about 0.3 points too optimistically, which accounted for more than half of its error on the deep benchmarks. The fix was to harvest exit descendants: for each training position, the legal moves that leave the category, rolled out in the orientation the router evaluates them. Forty thousand per category, added to the training set, removed the error mode. Warm-starting from Stage 9's own backgame network beat both a random start and a temporal-difference bootstrap.
Analysing where the trio's remaining error lay showed it was driven by the phase of the game rather than by which anchors were held: it climbed as the racer bore in, and was worst just after the holder hit a shot. Phase specialists trained on the matching slices of the same data did not beat the trio inside its region — but on the containment positions the plan-pair gate rejects (the racer's remaining block reading as a prime, so Stage 9 hands the position to a standard network), an early-containment specialist cut the 1-ply PR from 4.90 to 3.69 and halved the blunders. It is trained from the deep network on the early-containment rows of the trio's data and used only there.
The containment rule is new, so its region had almost no rolled-out data: about 1% of the existing piles. We played 3,000 games from the containment seeds with Stage 11 itself, excluding every benchmark position and candidate in both orientations, and kept the positions that satisfied the rule — 35,266 fresh boards, plus 9,734 already-rolled rows — and rolled them out at 1,296 paths with 3-ply play. A network trained on those 45,000 rows was excellent on the containment benchmark and worse than Stage 9 on the 237 unlimited-game positions the rule routes to it: those are ordinary-game containment positions (a late hit in a bear-off, a closeout) that never appeared in the family data. Adding the 291,000 rows of the general training corpus that satisfy the same rule, with the family rows weighted four times over, restored the unlimited-game slice to Stage 9's level and made the network better on the family benchmark too. That is the general lesson of this work: a specialist network must be trained on its whole region, never on the benchmark's corner of it.
The snake had no data anywhere — a hundred rows in a 1.8-million-row corpus — and could not be harvested from self-play: the model's play was so poor in these positions that it broke the prime or ran within a move or two, and 900 games yielded barely two snake decisions each. So the region was sampled directly. Five thousand synthetic snakes (prime length and placement, spare checkers spread from the far side to home, a crunched opponent with up to five checkers off and one to three stragglers) seeded short 2-ply games in which the holder preferred structure-keeping moves, and every decision contributed its position, its five best candidates and up to three release candidates — the moves the network must value correctly to know when to let go. That produced 98,000 distinct boards, of which 22,000 were rolled out at 648 paths with 3-ply play; snake games are long enough that a rollout costs six times a containment one. The network trains in minutes on that set, and it is the first Open Sage network with any competence in the region: its 1-ply error against the training targets fell from 0.30 to 0.054 of a point.
| Network | Training rows | Source | Warm start |
|---|---|---|---|
| Deep backgame | 245,000 + 40,000 exits | Stage 9/10 rolled piles, split by category; one-step exit descendants | Stage 9 backgame network |
| Middle backgame | 167,000 + 40,000 exits | as above | Stage 9 backgame network |
| Double anchor | 94,000 + 40,000 exits | as above | Stage 9 backgame network |
| Early containment | phase slice of the above | early-containment rows of the trio's data | Deep backgame network |
| Containment | 45,000 family + 291,000 general | Self-play from the containment seeds, rolled out; the rule's rows of the general corpus | Stage 9 priming-vs-racing network |
| Snake | 22,000 | Synthetic seeds + structure-keeping self-play, rolled out | Stage 9 priming-vs-racing network |
Every figure is a Performance Rating — lower is better — at the engine's own 1-ply, 2-ply and 3-ply evaluation and at its three truncated-rollout levels, 1T, 2T and 3T (the counterparts of XG's Roller, Roller+ and Roller++). Both engines are scored against the same references, with every pick of either engine rollout-graded. XG is compared against the engine on the bot performance page; on Paskogammon, where both engines have been run at every level, the comparison is below.
| Benchmark | Decisions | 1-ply | 2-ply | 3-ply | 1T · Roller | 2T · Roller+ | 3T · Roller++ | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| S9 | S11 | XG | S9 | S11 | XG | S9 | S11 | XG | S9 | S11 | XG | S9 | S11 | XG | S9 | S11 | XG | ||
| Unlimited game | 17,535 | 2.59 | 2.55 | — | 1.66 | 1.68 | — | 0.60 | 0.60 | 0.57 | 0.50 | 0.49 | 0.53 | 0.26 | 0.27 | 0.41 | 0.23 | 0.23 | 0.34 |
| Five-point match | 18,292 | 2.26 | 2.22 | — | 1.30 | 1.27 | — | 0.57 | 0.56 | 0.54 | 0.49 | 0.47 | 0.51 | 0.26 | 0.26 | 0.44 | 0.22 | 0.23 | 0.36 |
| Paskogammon | 2,556 | 7.74 | 6.14 | — | 5.42 | 4.57 | 4.38 | 3.46 | 2.28 | 2.97 | 2.67 | 1.65 | 2.63 | 2.21 | 1.46 | 2.44 | 1.47 | 0.88 | 1.92 |
| Backgame | Decisions | 1-ply | 2-ply | 3-ply | 1T · Roller | 2T · Roller+ | 3T · Roller++ | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| S9 | S11 | XG | S9 | S11 | XG | S9 | S11 | XG | S9 | S11 | XG | S9 | S11 | XG | S9 | S11 | XG | ||
| 2‑1 | 1,000 | 6.59 | 3.57 | — | 4.51 | 2.26 | 2.84 | 2.70 | 1.01 | 1.64 | 2.03 | 0.61 | 1.67 | 1.72 | 0.41 | 0.98 | 1.08 | 0.38 | 0.93 |
| 3‑1 | 1,001 | 4.44 | 3.28 | — | 2.77 | 2.22 | 2.22 | 0.93 | 0.88 | 0.95 | 0.92 | 0.66 | 0.91 | 0.66 | 0.34 | 0.59 | 0.53 | 0.32 | 0.48 |
| 3‑2 | 1,008 | 4.22 | 3.31 | — | 2.47 | 2.07 | 2.39 | 1.09 | 0.91 | 1.31 | 1.40 | 0.65 | 0.95 | 0.75 | 0.44 | 0.61 | 0.74 | 0.25 | 0.51 |
| 4‑1 | 1,001 | 3.70 | 2.74 | — | 1.76 | 1.77 | 1.83 | 0.85 | 0.78 | 0.82 | 1.18 | 0.55 | 0.69 | 0.52 | 0.24 | 0.47 | 0.56 | 0.19 | 0.45 |
| 4‑2 | 1,000 | 4.16 | 3.25 | — | 2.20 | 2.16 | 1.96 | 1.04 | 0.73 | 1.05 | 1.27 | 0.63 | 0.79 | 0.55 | 0.36 | 0.61 | 0.51 | 0.27 | 0.43 |
| 5‑1 | 1,003 | 3.36 | 2.56 | — | 1.69 | 1.32 | 1.77 | 0.95 | 0.66 | 0.76 | 1.20 | 0.49 | 0.60 | 0.62 | 0.26 | 0.41 | 0.57 | 0.23 | 0.32 |
| 5‑2 | 1,002 | 3.88 | 2.61 | — | 1.74 | 1.53 | 1.60 | 0.57 | 0.49 | 0.82 | 1.37 | 0.53 | 0.63 | 0.52 | 0.33 | 0.41 | 0.52 | 0.24 | 0.30 |
| 4‑3 | 1,002 | 4.56 | 3.24 | — | 2.04 | 1.69 | 2.13 | 0.81 | 0.74 | 0.80 | 1.23 | 0.52 | 0.76 | 0.72 | 0.24 | 0.55 | 0.54 | 0.19 | 0.44 |
| 5‑3 | 1,003 | 3.79 | 2.80 | — | 1.79 | 1.86 | 2.13 | 0.96 | 1.02 | 0.89 | 1.27 | 0.68 | 0.81 | 0.69 | 0.35 | 0.46 | 0.50 | 0.21 | 0.44 |
| 5‑4 | 1,001 | 5.97 | 3.10 | — | 2.84 | 2.45 | 2.74 | 1.42 | 0.93 | 1.15 | 1.53 | 0.76 | 0.97 | 0.70 | 0.47 | 0.55 | 0.63 | 0.30 | 0.49 |
| Pooled | 10,021 | 4.47 | 3.05 | — | 2.38 | 1.93 | 2.16 | 1.13 | 0.82 | 1.02 | 1.34 | 0.61 | 0.88 | 0.75 | 0.34 | 0.56 | 0.62 | 0.26 | 0.48 |
| Benchmark | Decisions | 1-ply | 2-ply | 3-ply | 1T · Roller | 2T · Roller+ | 3T · Roller++ | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| S9 | S11 | XG | S9 | S11 | XG | S9 | S11 | XG | S9 | S11 | XG | S9 | S11 | XG | S9 | S11 | XG | ||
| Containment | 3,270 | 9.65 | 6.82 | — | 9.63 | 4.73 | 14.62 | 5.09 | 2.54 | 11.25 | 3.78 | 2.29 | 7.76 | 3.20 | 1.30 | 7.85 | 2.53 | 1.05 | 6.16 |
| Massive backgame | 2,103 | 9.78 | 5.71 | — | 7.88 | 4.33 | 5.76 | 4.89 | 2.56 | 4.44 | 3.85 | 1.91 | 3.14 | 3.21 | 1.39 | 2.75 | 2.80 | 1.09 | 2.64 |
| Snake | 978 | 64.87 | 21.08 | — | 39.17 | 21.00 | 39.92 | 26.51 | 21.42 | 36.64 | 28.25 | 11.19 | 32.99 | 20.96 | 8.42 | 32.38 | 14.79 | 5.85 | 32.24 |
Paskogammon is where the backgame work shows up in whole games: played from its
scattered opening position, it produces far more backgames and containment games than
standard backgammon, so it is the one full-game benchmark on which these networks fire
often enough to move the headline number. It is also the one on which XG has been run at
every level it offers, from 2-ply to Roller++ — the games are exported to XG as native
.xg archives, which carry the non-standard starting position explicitly, and
batch-analysed once per level. The table scores Stage 9, Stage 11 and XG against the same
rollout reference over the benchmark's 2,556 decisions, grouped by matched level; the
best of each group is highlighted.
| Engine | PR | Checker | Cube | Pure race | Racing | Attacking | Priming | Anchoring |
|---|---|---|---|---|---|---|---|---|
| Stage 9 3T | 1.47 | 1.36 | 2.21 | 0.08 | 1.46 | 1.11 | 0.80 | 2.45 |
| Stage 11 3T | 0.88 | 0.85 | 1.07 | 0.08 | 0.59 | 0.96 | 0.44 | 1.71 |
| XG Roller++ | 1.92 | 1.99 | 1.37 | 0.01 | 1.36 | 2.35 | 1.10 | 3.39 |
| Stage 9 2T | 2.21 | 2.27 | 1.80 | 0.03 | 2.21 | 2.25 | 0.90 | 3.54 |
| Stage 11 2T | 1.46 | 1.51 | 1.08 | 0.03 | 1.36 | 1.79 | 0.51 | 2.35 |
| XG Roller+ | 2.44 | 2.59 | 1.45 | 0.02 | 1.97 | 2.48 | 1.54 | 4.19 |
| Stage 9 1T | 2.67 | 2.60 | 3.16 | 0.08 | 2.55 | 1.74 | 2.17 | 4.24 |
| Stage 11 1T | 1.65 | 1.64 | 1.76 | 0.08 | 1.14 | 1.37 | 1.02 | 3.31 |
| XG Roller | 2.63 | 2.69 | 2.23 | 0.05 | 2.36 | 2.86 | 1.54 | 4.15 |
| Stage 9 4P | 2.68 | 2.75 | 2.20 | 0.14 | 2.51 | 2.46 | 1.52 | 4.37 |
| Stage 11 4P | 1.90 | 1.92 | 1.72 | 0.14 | 1.64 | 2.43 | 1.04 | 2.88 |
| XG 4-ply | 2.62 | 2.66 | 2.32 | 0.05 | 2.22 | 2.76 | 1.58 | 4.32 |
| Stage 9 3P | 3.46 | 3.53 | 2.98 | 0.03 | 3.50 | 3.62 | 2.20 | 4.84 |
| Stage 11 3P | 2.28 | 2.35 | 1.78 | 0.03 | 1.64 | 3.69 | 1.21 | 3.53 |
| XG 3-ply | 2.97 | 2.96 | 3.04 | 0.05 | 2.55 | 3.21 | 1.77 | 4.81 |
| Stage 9 2P | 5.42 | 5.39 | 5.57 | 0.16 | 5.02 | 5.19 | 4.11 | 8.02 |
| Stage 11 2P | 4.57 | 4.37 | 5.93 | 0.16 | 3.97 | 5.16 | 3.95 | 6.30 |
| XG 2-ply | 4.38 | 4.50 | 3.58 | 0.23 | 3.52 | 5.83 | 2.98 | 6.50 |
| Stage 9 1P | 7.74 | 7.67 | 8.21 | 0.26 | 6.94 | 8.51 | 6.56 | 10.61 |
| Stage 11 1P | 6.14 | 5.93 | 7.58 | 0.26 | 5.18 | 8.55 | 4.70 | 8.18 |
The engine, the benchmarks, the seed positions, the references and every training
script are in the open-source
bgsage
repository (MPL-2.0). The backgame benchmarks live under
backgame_ref_positions/benchmark/ — one starting.txt of seeds, one
benchmark.txt of decisions and one rollout.jsonl reference per
folder — and are generated by scripts/backgame_benchmark.py, seeded for the
three families by scripts/build_family_seeds.py, and scored by
scripts/score_backgame_pr.py, which reports the PR alongside the share of
picks that are rollout-graded. The routing rules are in cpp/src/neural_net.cpp
with Python reference implementations in scripts/containment_rule.py and
scripts/snake_rule.py, and the model itself is the stage11s entry
of the weight registry. The training-data scripts are named in the table above by their
source; the rollouts they need are distributed by orchestrators that are not part of the
public repository, but every position file and every rolled-out target is.
| Cubeless | Win | Gammon | Backgammon |
|---|---|---|---|
| Player | |||
| Opponent |
| Cube action | Equity |
|---|---|
| No double | |
| Double / Take | |
| Double / Pass |