Engine Study · Bot Performance

Benchmarking Open Sage against eXtreme Gammon

A quantitative comparison of the Open Sage evaluation engine — which powers Backgammon Sage Pro — against eXtreme Gammon's bot engine (XG).

Abstract

We compare the Open Sage bot engine against eXtreme Gammon (XG) across matched evaluation levels — direct neural-network evaluation, multi-ply lookahead, and truncated rollouts. Because XG exposes no programmatic interface, we score it through its Batch Analysis feature. We use three complementary methods. The first constructs a layered set of ground-truth evaluations — trusting fast evaluations where the choice is clear and escalating close calls to full rollouts — and scores each engine as a Performance Rating (PR) against that reference, across both 500 money games (17,535 decisions) and 130 five-point matches (18,292 decisions). For the money games we then rebuild that entire comparison against XG's own tiered analysis as the reference — the exact mirror. The second method isolates the money positions on which Sage 3T and XG Roller++ genuinely disagree and adjudicates each against both engines' full rollouts. The same reference-and-score method is applied to thirteen benchmarks of backgames, containment games and the snake — the positions engines have historically played worst. Across the strength studies the two engines are close: scored against Sage's reference, Sage is ahead at the matched truncated-rollout levels most users rely on; scored against XG's, XG is ahead or level. The third method turns the question around: on 290 real tournament matches already analyzed in XG, we re-analyze each in Sage and compare the Performance Rating the two engines assign to every player. The two ratings track each other very closely: across 580 player ratings they have a 98% correlation, and the average difference between them is +0.002 PR — statistically indistinguishable from zero.

01
Introduction

Why measure against XG?

eXtreme Gammon (XG) is a standard reference for computer backgammon analysis, and its evaluation engine is widely considered one of the strongest available. Any new engine that aims to be taken seriously has to answer a simple question: how does it compare to XG? This study sets out to answer that quantitatively for Open Sage, the neural-network engine behind Backgammon Sage Pro.

Our goal was to compare Open Sage's evaluations against XG's at comparable levels of effort. The natural experiment — having the engines play many millions of games head-to-head and tallying the result — is not practical: the XG desktop app exposes no programmatic interface, so feeding positions and moves between the two engines is a manual, click-through process. Reaching a sample size large enough to resolve the small differences between two strong engines would take far too long.

Instead, we use XG's Batch Analysis feature, which can score hundreds of transcribed games in a single unattended pass. Open Sage plays both sides of many games; each game is exported in a standard text format that XG can import; XG analyzes the lot; and we parse its verdicts back out. On top of that shared foundation we apply two distinct methods, described in Section 3 and Section 4.

A third study (Section 5) sets the simulations aside altogether and works from real tournament matches analyzed in XG. Its purpose is different: rather than asking which engine is stronger, it asks whether a player who analyzes a match in Sage gets the same Performance Rating that XG would give — so that years of intuition about what a PR means in XG carry straight over to Sage.

02
Background

Levels of evaluation

Both engines can evaluate a position at a range of depths, trading accuracy for speed. Understanding these levels is essential to reading the results, because the comparison is always between matched levels of the two engines.

Direct evaluation (1-ply). The raw output of the neural network for a position, with no lookahead. This is the fastest setting and the foundation every deeper level is built on. (We follow XG's convention, in which 1-ply denotes the raw network evaluation.)

Multi-ply lookahead (2-, 3-, 4-ply). The engine looks several turns into the future, considering the opponent's replies and averaging over the possible dice rolls at each step, with a raw network evaluation at the end of each line. Each additional ply is more accurate and substantially more expensive.

Truncated rollouts (1T, 2T, 3T). Rather than searching to a fixed depth, the engine plays out many short simulated games, truncates them after a few turns, and evaluates the resulting position with multi-ply lookahead. Variance reduction cancels out much of the luck in the sampled dice. These are XG's "Roller" settings, and they are stronger than fixed-depth search while remaining far cheaper than a full rollout.

Full rollout. Many simulated games played all the way to completion. Run at sufficient volume with variance reduction, a full rollout is the closest thing to ground truth that a backgammon engine can produce, and we use it as the reference standard throughout this study.

Matched levels: Sage and XG

Family Sage XG equivalent What it does
Direct 1-ply 1-ply (raw eval) A single neural-network evaluation of the position — no lookahead.
Multi-ply 2P · 3P · 4P 2-ply · 3-ply · 4-ply Searches the opponent's replies several turns deep, averaging over the dice.
Truncated rollout 1T · 2T · 3T Roller · Roller+ · Roller++ Many short simulated games, truncated after 5–7 turns and evaluated with multi-ply, with variance reduction. 3T/Roller++ use 3-ply decisions; 2T/Roller+ use 2-ply; 1T/Roller use 1-ply.
Full rollout Rollout Rollout Simulated games played to completion — the reference "truth" in this study.
Performance Rating (PR). Throughout, an engine's accuracy is summarized as PR — the average equity error per decision, multiplied by 500, measured against the rolled-out reference. A PR of 0 means perfect agreement with the truth; lower is better.
03
Method One

Rollout PR

The first method scores every engine against a fixed set of decisions whose "true" best play we have established as accurately as possible. It follows the same approach used in the well-known 2012 XG study that compared XG against a number of other bots.

We began by simulating 500 money games of Sage 3P playing itself. A moderately strong level like 3P produces a realistic distribution of positions across every game plan — racing, attacking, priming, anchoring — which is exactly the variety we want to test against. We then ran the whole study a second time for 130 five-point matches, where the score on the board changes the value of every decision; results for both appear below.

Building the ground truth

Establishing the true best play for every decision by full rollout would be enormously expensive — and unnecessary, because most decisions are not close. We therefore build the reference in three escalating passes, spending rollout effort only where the choice is genuinely in doubt. The same recipe is applied to both data sets; each pass settles the decisions it can resolve confidently and hands the rest down:

Pass 1 · Sage 3P

Trust 3-ply when the choice is clear

Every decision is evaluated at 3-ply. Where the best move beats the next-best by more than 0.05 equity, the 3P verdict is accepted as truth and the decision is settled.

0.05
equity gap to settle
Closer calls (within 0.05) pass down to Pass 2
Pass 2 · Sage 3T

Re-check close calls with a truncated rollout

The remaining decisions are re-evaluated at 3T. Where the margin is now wider than 0.02 equity, the 3T verdict is accepted as truth.

0.02
equity gap to settle
Anything still within 0.02 passes down to Pass 3
Pass 3 · Full rollout

Roll out the hardest decisions

Everything still inside 0.02 is settled by a full Sage rollout: 3P decisions for both checker play and cube actions, run in batches of 1,296 paths until the 95% confidence band on the equity falls under 0.005 — or a ceiling of 20,736 paths (16 × 1,296) is reached. The slowest backgame positions took well over an hour each.

0.005
target 95% band

The result is a layered reference in which every decision is resolved at exactly the depth its difficulty demands. We ran this same three-pass recipe over both data sets — the 500 money games and the 130 five-point matches — and the resulting tier compositions are shown alongside each set's results below.

Scoring an engine

With a layered set of trusted evaluations in hand, scoring any engine is straightforward: we present it with each benchmark decision, record the equity error of the move or cube action it chooses relative to the truth, and report the average error × 500 as a PR. The same total is broken out into checker-play and cube PRs, and further by game plan.

Scoring XG required the extra step described in the introduction. Because XG has no API, we ran its Batch Analysis over the 500 money-game and 130 match transcripts using a custom level — 3-ply decisions, upgrading to the level under test wherever it disagreed — and parsed XG's preferred decision for each position out of the .xg files it produces, scoring those against the same reference equities.

Money games — ground-truth composition

Across the 500 money games, the three passes resolved 16,889 positions (17,535 decisions) as follows:

Settled at 3P  7,652 Settled at 3T  3,260 Rolled out  5,977 Money games  16,889 positions · 17,535 decisions

Money games — head to head

The five matched levels, scored over those 17,535 decisions. Lower PR is better; the stronger engine in each pair is highlighted.

Sage 3T
Truncated rollout · 3-ply decisions · 7-turn truncation
0.23
vs
0.34
XG Roller++
Truncated rollout · 3-ply decisions · 7-turn truncation
Sage 3T leads by 0.11 PR
Sage 2T
Truncated rollout · 2-ply decisions
0.27
vs
0.41
XG Roller+
Truncated rollout · 2-ply decisions
Sage 2T leads by 0.14 PR
Sage 1T
Truncated rollout · 1-ply decisions · 5-turn truncation · 72 paths
0.49
vs
0.53
XG Roller
Truncated rollout · 1-ply decisions · 5-turn truncation · 42 paths
Sage 1T leads by 0.04 PR
Sage 4P
4-ply lookahead search
0.43
vs
0.47
XG 4-ply
4-ply lookahead search
Sage 4P leads by 0.04 PR
Sage 3P
3-ply lookahead search
0.60
vs
0.57
XG 3-ply
3-ply lookahead search
Effectively tied — XG 3-ply ahead by 0.03 PR

Money games — full breakdown

PR for every engine and level tested, broken out by decision type and game plan. Lower is better.

Engine PR Checker Cube Pure Race Racing Attacking Priming Anchoring
Sage 3T 0.230.200.410.020.250.200.320.32
XG Roller++ 0.340.330.380.040.440.250.410.48
Sage 2T 0.270.240.430.020.320.200.400.36
XG Roller+ 0.410.410.390.050.600.310.470.54
Sage 1T 0.490.500.430.040.580.430.620.64
XG Roller 0.530.540.480.050.640.450.710.67
Sage 4P 0.430.420.520.080.540.390.450.60
XG 4-ply 0.470.460.520.060.590.400.570.59
Sage 3P 0.600.600.610.130.780.530.670.75
XG 3-ply 0.570.570.580.050.720.480.730.72
Sage 2P 1.681.442.910.401.841.861.871.77
Sage 1P 2.552.423.240.422.712.793.132.74
Open Sage eXtreme Gammon Sage 2P and 1P have no XG counterpart in this run and are shown for reference.
Sage's evaluations are stronger than the equivalent XG evaluation at every matched level except 3-ply, where XG edges ahead by 0.03 PR — a difference inside the noise. The two engines are close, and at the truncated-rollout levels that most users rely on Sage holds a clear edge: 0.23 against Roller++'s 0.34, and 0.27 against Roller+'s 0.41.

Money games — scored against XG's own reference

Every number above grades the engines against Sage's tiered reference. The obvious objection is home-field advantage — Sage is being measured against its own rollouts. So we built the exact mirror: the same 500 games and the same decisions, but with XG's own analysis as the truth at every tier. XG 3-ply settles the 3P-tier decisions, XG Roller++ the 3T-tier decisions, and XG's own full rollout the rolled-out decisions — XG's tier-for-tier analogue of the Sage reference. Every engine is then re-scored against it.

Engine PR Checker Cube Pure Race Racing Attacking Priming Anchoring
Sage 3T 0.280.260.410.020.400.230.310.37
XG Roller++ 0.210.200.230.010.300.180.220.25
Sage 2T 0.300.290.360.020.390.280.360.38
XG Roller+ 0.300.320.230.010.480.250.260.40
Sage 1T 0.460.470.420.030.610.420.500.57
XG Roller 0.400.410.350.030.530.340.530.45
Sage 4P 0.390.380.430.060.550.360.420.42
XG 4-ply 0.340.330.400.030.460.330.410.36
Sage 3P 0.500.490.560.080.670.470.560.56
XG 3-ply 0.410.410.400.030.540.360.520.46
Sage 2P 1.381.172.440.301.571.501.551.45
Sage 1P 2.192.072.810.362.272.442.712.30
Open Sage eXtreme Gammon Reference: XG 3-ply (3P tier) · XG Roller++ (3T tier) · XG full rollout (rollout tier).
Against XG's own reference the ranking turns around: XG is ahead at every matched level except 2T, where Sage 2T and XG Roller+ are level at 0.30, and XG Roller++ scores 0.21 against Sage 3T's 0.28. Both tables carry the same built-in lean. Nearly two-thirds of the decisions are settled by the reference's own 3-ply or truncated rollout, and the level that settles a tier scores almost nothing on it — XG 3-ply and XG Roller++ here, Sage 3P and Sage 3T in the table above. The rolled-out decisions are the fairer test, because there a full rollout rather than a level under test is the judge: XG's rollout puts Roller++ narrowly ahead of Sage 3T (0.56 against 0.60), while Sage's rollout puts Sage 3T well ahead (0.51 against 0.76).
Two references, two verdicts. By Sage's reference Sage 3T leads XG Roller++ 0.23 to 0.34; by XG's, Roller++ leads 0.21 to 0.28. At the top level each engine comes out ahead against its own analysis — the same pattern the disputed-position study (Section 4) finds where the two engines disagree outright. The two engines are very close, and which one leads depends on whose analysis is taken as truth.

Match play — ground-truth composition

The match set runs the same three passes, with one addition: each player's away-score and the Crawford flag are threaded through every evaluation, so the truth is computed in match-equity (MWC) space against the correct score. Across the 130 five-point matches, the passes resolved 17,892 positions (18,292 decisions) as follows:

Settled at 3P  7,522 Settled at 3T  3,460 Rolled out  6,910 Match play  17,892 positions · 18,292 decisions

Match play — head to head

The five matched levels, scored over those 18,292 decisions. Lower PR is better; the stronger engine in each pair is highlighted.

Sage 3T
Truncated rollout · 3-ply decisions · 7-turn truncation
0.23
vs
0.36
XG Roller++
Truncated rollout · 3-ply decisions · 7-turn truncation
Sage 3T leads by 0.13 PR
Sage 2T
Truncated rollout · 2-ply decisions
0.26
vs
0.44
XG Roller+
Truncated rollout · 2-ply decisions
Sage 2T leads by 0.18 PR
Sage 1T
Truncated rollout · 1-ply decisions · 5-turn truncation · 72 paths
0.47
vs
0.51
XG Roller
Truncated rollout · 1-ply decisions · 5-turn truncation · 42 paths
Sage 1T leads by 0.04 PR
Sage 4P
4-ply lookahead search
0.41
vs
0.46
XG 4-ply
4-ply lookahead search
Sage 4P leads by 0.05 PR
Sage 3P
3-ply lookahead search
0.56
vs
0.54
XG 3-ply
3-ply lookahead search
Effectively tied — XG 3-ply ahead by 0.02 PR

Match play — full breakdown

PR for every engine and level tested on the 5-point match set, broken out by decision type and game plan. Lower is better.

Engine PR Checker Cube Pure Race Racing Attacking Priming Anchoring
Sage 3T 0.230.200.460.070.220.200.280.31
XG Roller++ 0.360.350.440.080.440.320.330.49
Sage 2T 0.260.240.400.020.350.180.310.35
XG Roller+ 0.440.440.440.090.540.440.390.55
Sage 1T 0.470.470.470.060.550.450.550.56
XG Roller 0.510.510.560.090.630.490.500.65
Sage 4P 0.410.410.470.050.500.370.440.53
XG 4-ply 0.460.450.580.090.590.440.420.60
Sage 3P 0.560.540.710.050.730.550.590.63
XG 3-ply 0.540.530.620.090.680.530.520.68
Sage 2P 1.271.231.580.391.461.241.491.39
Sage 1P 2.222.212.350.382.432.412.442.47
Open Sage eXtreme Gammon Sage 2P and 1P have no XG counterpart in this run and are shown for reference.
The match results echo the money study: Sage is stronger than the equivalent XG evaluation at every matched level except 3-ply, where XG edges ahead by 0.02 PR — well inside the noise. Sage's clearest edge is again at the truncated-rollout levels (3T and 2T) that most users rely on: 0.23 against Roller++'s 0.36 and 0.26 against Roller+'s 0.44.

Backgames, containment games and the snake

Ordinary self-play games rarely produce a deep backgame, a containment game or a far-side prime holding a single straggler, so the money and match sets say little about how an engine plays them — and they are the positions where engines have historically been weakest. Thirteen position-family benchmarks measure them directly, each built like the money benchmark — real decisions, each with a rollout-grade reference — but with the decisions drawn from inside the family:

  • The ten classic backgames, named by the two points the backgame player holds in the opponent's home board — the 2‑1 backgame holds the opponent's 2- and 1-points, the 5‑4 the 5- and 4-points, and so on. Each holds about 1,000 decisions, recorded only while both anchors are still held and the holder is behind in the race.
  • Containment games, defined by the escaper: one side has borne off some checkers and has one to three checkers that were hit and must run the whole board home, while the other side, with a lost race, arranges whatever it has left to keep hitting them.
  • Massive backgames — three or more anchors in the opponent's home board, or two anchors with seven or more checkers back.
  • The snake — a run of four or more consecutive made points on the opponent's half of the board trapping a single straggler while the opponent's other checkers are crunched home. A priming game rather than a backgame, and one of the hardest shapes in backgammon to play.

Each family starts from a small set of seed positions. Sage plays unlimited games out of those seeds against itself at 3-ply, every decision that arises while the position still belongs to the family is recorded, and every recorded decision is rolled out to produce its reference — 5,184 paths played to completion, with 3-ply checker and cube decisions for the first three half-moves and 2-ply thereafter, variance reduction on. Candidate-completion passes rolled out every move an engine under test actually chose, so each score below is against a rollout-graded candidate. A cube position where the opponent owns the cube is not a decision for the player on roll and is not scored.

Benchmark Decisions 2-ply 3-ply 4-ply 1T / Roller 2T / Roller+ 3T / Roller++
SageXG SageXG SageXG SageXG SageXG SageXG
2-1 backgame1,0002.262.841.011.641.031.220.611.670.410.980.380.93
3-1 backgame1,0012.222.220.880.950.670.720.660.910.340.590.320.48
3-2 backgame1,0082.072.390.911.310.780.990.650.950.440.610.250.51
4-1 backgame1,0011.771.830.780.820.600.530.550.690.240.470.190.45
4-2 backgame1,0002.161.960.731.050.690.550.630.790.360.610.270.43
5-1 backgame1,0031.321.770.660.760.460.640.490.600.260.410.230.32
5-2 backgame1,0021.531.600.490.820.480.630.530.630.330.410.240.30
4-3 backgame1,0021.692.130.740.800.550.580.520.760.240.550.190.44
5-3 backgame1,0031.862.131.020.890.580.630.680.810.350.460.210.44
5-4 backgame1,0012.452.740.931.150.730.780.760.970.470.550.300.49
Containment3,2704.7314.622.5411.252.0510.372.297.761.307.851.056.16
Snake97821.0039.9221.4236.6423.5336.1711.1932.998.4232.385.8532.24
Massive backgame2,1034.335.762.564.442.173.681.913.141.392.751.092.64
Open Sage eXtreme Gammon PR against each family's rollout reference; lower is better; the better engine of each pair is highlighted.
Pooled over the ten classic backgames (10,021 decisions), Sage's PR runs 1.93 at 2-ply, 0.82 at 3-ply and 0.26 at 3T, with blunders falling from 28 to none at all. Containment games and massive backgames are three to four times harder at every level — 1.05 and 1.09 at 3T — and the snake is in a class of its own: 21.00 at 2-ply and still 5.85 at 3T. It is also the one family where added ply does not reliably help while the truncated-rollout levels do, which is what a family of hold-or-release decisions looks like: the choice turns on how a long containment plays out, and only simulation gets at that. Read the snake row as a measure of how far the position family is from solved rather than as a ranking of Sage's levels.
XG on these benchmarks. Sage leads on all thirteen families and at every level pooled: on the ten backgames 1.93 against 2.16 at 2-ply and 0.26 against 0.48 at 3T. The margin widens with the difficulty of the family — at 3T, containment 1.05 against 6.16, massive backgames 1.09 against 2.64, and the snake 5.85 against 32.24. Only four of the 78 individual cells go the other way, all single backgames at 2-ply to 4-ply where both engines are already under 1.1. XG has no programmatic interface, so its decisions are replayed through Batch Analyze from native .xg exports (see the Appendix).
04
Method Two

Disputed positions

The Rollout PR method scores both engines against a shared reference. A sharper, more direct question is this: in realistic play, where do the two engines actually disagree on the best move — and when they do, which one is right?

We answer it on the hardest money decisions from Method One — the rolled-out positions, each of which now carries both a full Sage rollout and a full XG rollout. Among them we find every decision where Sage 3T and XG Roller++ chose differently, and ask which choice each rollout preferred. Having both rollouts is the whole point: every disagreement is judged against Sage's rollout and XG's own, so the answer never rests on trusting a single engine's truth.

  • Start from the rolled-out money benchmark — the hardest decisions, each settled by a full Sage rollout.
  • Have XG produce its own full rollout of the same positions (its Batch Rollout) — an independent second ground truth.
  • Keep the decisions where Sage 3T and XG Roller++ disagree, and score each engine's pick against each rollout's best move or cube action.
  • Report on the common set — disagreements where both engines' picks were rolled by both rollouts — so the two comparisons cover the identical positions.

Results

Across the rolled-out money positions — the hardest decisions in the study — the two engines disagree on 1,286 checker plays and 80 cube decisions. Each panel shows how often each engine matched the rollout's best decision — scored against XG's own rollout and against Sage's rollout in turn.

Checker play

Where Sage 3T and XG Roller++ chose different moves
5,687
rolled-out checker decisions
1,286
disagreements
1,139
scored (common set)
Matched the best move — by XG's own rollout
Avg equity error — Sage 3T 0.0030 · XG Roller++ 0.0030  (PR 1.51 vs 1.49)
Matched the best move — by Sage's rollout
Avg equity error — Sage 3T 0.0023 · XG Roller++ 0.0040  (PR 1.14 vs 1.98)
Sage 3T XG Roller++ Neither

Cube decisions

Where Sage 3T and XG Roller++ chose different cube actions
282
rolled-out cube decisions
80
disagreements
Matched the best action — by XG's own rollout
Avg equity error — Sage 3T 0.0110 · XG Roller++ 0.0074  (PR 5.50 vs 3.69)
Matched the best action — by Sage's rollout
Avg equity error — Sage 3T 0.0058 · XG Roller++ 0.0102  (PR 2.88 vs 5.08)
Sage 3T XG Roller++
Only 80 cube disagreements in all — a small sample, so read this panel as suggestive rather than decisive.
On checker play — the large majority of disagreements — each rollout sides with its own engine on which move is best: by XG's rollout XG Roller++ matches the best move more often (52.2% vs 37.7%), by Sage's rollout Sage 3T does (50.0% vs 38.1%). Measured by how much equity a disputed pick gives away, the two are level by XG's rollout (0.0030 each) and Sage 3T is closer by Sage's (0.0023 vs 0.0040). Cube disagreements are far rarer and mixed: Sage 3T matches the rolled-out cube action more often under both rollouts, but by XG's rollout its misses are the more expensive ones.
Why two rollouts matter. A single-reference result always invites the question "whose truth?" Here the two rollouts come from independent engines, and on the positions where two strong engines differ they differ too, each leaning toward the engine that ran it. The disputed positions do not separate Sage 3T and XG Roller++.
05
Method Three

Agreement on real matches

The first two methods ask which engine is stronger. A third question is just as important to anyone who uses an engine to study their own play: if you analyze a real match in XG, note your Performance Rating, then analyze the very same match in Sage — how close are the two ratings? A player who has spent years building intuition for what a given PR means in XG should get essentially the same number from Sage.

To test this directly, we took a large collection of real tournament matches that had already been analyzed in XG, re-analyzed every one of them from scratch in Sage, and compared the Performance Rating each engine assigned to each player.

The matches

The match files come from three 2026 tournaments, all 7-point matches, generously provided by Máté Fehér — already analyzed in eXtreme Gammon, exactly as a competitive player would study their own games.

UBC Texas 2026
100
matches analyzed
UBC Istanbul 2026
146
matches analyzed
UBC Japan 2026
44
matches analyzed

290 matches and 580 individual player ratings in all. One further match was set aside as a corrupted transcription.

Matched evaluation settings

Each match was re-analyzed in Sage at a 3-ply base, with an expert 3T pass — a 360-path truncated rollout — applied to the decisions where the player's actual move disagreed with the 3-ply best. This mirrors a strong XG analysis (a base ply for the clear decisions, escalating to a truncated rollout for the close ones), at matched strength: Sage 3P ≈ XG 3-ply and Sage 3T ≈ XG Roller++ (see Section 2). Each engine then computes a PR per player from its own evaluations — the number a user sees in each app.

Results

Pooling all 580 player ratings, the two engines agree almost exactly. The average difference is statistically indistinguishable from zero, and the spread of that difference is small next to the spread in PR itself.

+0.002
Mean difference (PR)
Statistically indistinguishable from zero — 95% confidence interval ±0.03, p = 0.90.
0.37
Std. dev of the difference
Small next to the 2.08 spread in PR itself — the disagreement on any one rating is minor relative to how much PR varies.
0.98
Rating correlation (r)
Sage and XG rate the same players almost identically.
Per-player Performance Rating XG Sage DifferenceSage − XG
Average 4.36 4.36 +0.002
Standard deviation 2.08 2.10 0.37
95% range 1.52 – 9.36 1.44 – 9.67 −0.76 – +0.74
In practical terms, a player who analyzes a match in Sage will — in the large majority of cases — see essentially the same Performance Rating that XG would give. As a measure of how well a match was played, the two engines are interchangeable.
06
Conclusion

Two strong engines, very close

Across the strength studies, Open Sage 3T and XG Roller++ are two of the strongest backgammon evaluations available, and close to each other. The Rollout PR study — run for both money play and 5-point matches — finds Sage stronger at every matched level except 3-ply, where the two engines are within noise of each other, with the clearest edge at the truncated-rollout levels. When the money-game study is rebuilt with XG's own analysis as the reference, the ranking turns around: XG is ahead at every matched level except 2T, where the two are level. The disputed-position study, which looks at positions where the two engines disagree at the strongest evaluation level and scores them against both engines' rollouts, is level on checker play — each rollout favouring the engine that ran it — with cube disagreements too rare and too split to call.

The honest summary, then, is that the two engines play at a very similar, very high level, and that each scores best against its own analysis: by Sage's reference its truncated rollouts get right decisions that XG's Roller levels miss, by XG's the reverse holds, and on the positions where the two engines disagree outright neither engine's rollout puts the other clearly behind. The family benchmarks — backgames, containment games and the snake — are where the two engines separate: measured to the same standard and at every matched level, Sage leads on all thirteen, and by a factor of three or more on containment games and the snake.

And the third study answers the question a Sage user actually faces: across 290 real tournament matches already rated by XG, Sage produces essentially the same Performance Rating — the average difference between a player's Sage and XG rating is statistically indistinguishable from zero, and the spread of that difference is small next to the spread in PR itself. Whether the test is strength against a rolled-out truth or simple agreement on how a real game was played, Open Sage and XG land in the same place.

Try Backgammon Sage Pro Back to home
A
Appendix

Reproducing these results

Open Sage is released as open source, and the complete pipeline behind both methods ships in the engine's repository. Everything below runs against the bgsage package and a local build of its engine alone — no external services, datasets, or infrastructure are required. The only manual dependency is eXtreme Gammon itself, used to regenerate the XG columns, since XG has no programmatic interface.

github.com/markbgsage/bgsage Open Sage engine · MPL-2.0 · run all scripts from the repo root
Prerequisites. Build the bgsage engine (the bgbot_cpp extension) following the repository's instructions for your platform. Reproducing any XG column additionally requires eXtreme Gammon (Windows) for its Batch Analysis step. Every script resolves its paths inside the repository, so run them from the bgsage/ root.

Reproducing it with Claude Code

You don't have to drive the pipeline by hand. Claude Code — Anthropic's agentic command-line coding tool — can build the engine and run the benchmarks for you. Point it at a checkout of the repository and describe what you want; it works through the same steps detailed below:

›"Build this engine, then score Sage 3T and Sage 3P against the shipped benchmark and show me the PR, checker PR, and cube PR for each."
›"Reproduce the Rollout PR table: score every Sage level — 1-ply through 4-ply, plus 1T, 2T and 3T — against data/money_benchmark/benchmark.json.gz and print it as a table I can compare to the study."
›"Rebuild the ground-truth benchmark from scratch with the three build passes, then score Sage 3T against it."
›"I have my own bot in ./mybot — for each decision in the shipped benchmark, evaluate it with my bot and compute its PR against the reference equities."

Method One — Rollout PR

The benchmark — the layered set of trusted evaluations — is built and scored by a single script, scripts/benchmark_money.py. There are two ways in.

A

Score against the shipped dataset

No rebuild · validate the table or score a new engine

The repository ships the assembled benchmark as data/money_benchmark/benchmark.json.gz (the uncompressed JSON exceeds GitHub's file-size limit; the scripts read the gzip transparently). Scoring any Sage level reproduces its row in the Section 3 table:

# lower PR is better
python scripts/benchmark_money.py score --level 3ply        # Sage 3P
python scripts/benchmark_money.py score --level truncated3  # Sage 3T
# --level: 1ply 2ply 3ply 4ply | truncated1/2/3 | rollout

To score a different engine entirely, apply the same rule the script uses internally: for every decision in the dataset, take your engine's choice and compare its equity to the stored reference ("rollout") equity — the mean error × 500 is its PR.

B

Rebuild the ground truth from scratch

Re-derive every decision · seeded, so it reproduces

Three adaptive passes re-create the dataset. The games are seeded (500 games from seed 1), so the decisions and reference equities come out the same:

# 1. simulate 500 Sage-3P self-play games (+ XG transcripts)
python scripts/benchmark_money.py build --stages pass1 --n-games 500 --workers 6
# 2. re-evaluate every decision within 0.05 at 3T
python scripts/benchmark_money.py build --stages pass2 --n-threads 16
# 3. roll out every decision still within 0.02
python scripts/benchmark_money.py build --stages pass3 --n-threads 16

Pass 3 writes data/money_benchmark/benchmark.json (and its .gz); score it exactly as in Option A. Omitting --stages runs all three passes in order.

Reproducing the XG columns

XG has no API, so its results come from XG's Batch Analysis. Pass 1 writes XG-import transcripts to data/money_benchmark/xg/ — these are not shipped, so regenerate them with pass 1 if you started from the dataset alone.

  1. In XG, Batch-Analyze the xg/ folder with a custom level — 3-ply decisions, upgrading to the level under test on disagreements — with Save Games after analyze checked. XG writes one .xg per game.
  2. Run python scripts/benchmark_pr_xg_levels_all.py --benchmark money, which reads XG's top decision per position at every analysed level from the .xg files and scores it against the same reference equities, printing the same PR breakdown and recording XG's picks beside it.

The XG-reference money table in Section 3 — and the disputed-position study — instead need XG's rollouts of the hardest positions. Run XG's Batch Rollout on those positions (Save Games checked); python scripts/xg_benchmark_report.py parses them. The XG-reference table then re-scores every level from the picks each scorer recorded. See Method Two below.

# XG-reference table: harvest XG Roller++ (the 3T tier), then re-score every level;
# Sage's picks come from benchmark_money.py score --level <level> --model stage11
python scripts/xg_harvest_results.py --benchmark money --set rollerpp
python scripts/benchmark_pr_xg_reference_all.py --sage-suffix stage11

Match play

The 5-point match study is reproduced the same way with scripts/benchmark_match.py, the match-play twin of benchmark_money.py — the match length and number of matches are arguments, and the match score is threaded through every evaluation. Its ground truth ships as data/match_benchmark/5pt/benchmark.json.gz.

# score against the shipped match dataset (no rebuild)
python scripts/benchmark_match.py score --match-length 5 --level truncated3
# or rebuild: 130 seeded 5-pt matches, three adaptive passes
python scripts/benchmark_match.py build --match-length 5 --n-matches 130 --stages pass1 --workers 6
python scripts/benchmark_match.py build --match-length 5 --n-matches 130 --stages pass2 --n-threads 16
python scripts/benchmark_match.py build --match-length 5 --n-matches 130 --stages pass3 --n-threads 16
# XG columns: batch-analyze data/match_benchmark/5pt/xg/ per level, then
python scripts/benchmark_pr_xg_levels_all.py --benchmark match --match-length 5

Method Two — Disputed Positions

The disagreement study draws from the same money benchmark built for Method One, cross-referenced against XG's own full rollout of the hardest positions. The only manual dependency is XG's Batch Rollout; a single script does the rest.

Money games

  1. Build the money benchmark (Method One, Option B). Pass 1 also writes the XG-import transcripts to data/money_benchmark/xg/.
  2. In XG, Batch-Rollout those positions with Save Games after analyze checked, so each rolled-out decision carries XG's own rollout equities in the resulting .xg files.
  3. Harvest and report:
python scripts/xg_benchmark_report.py
python scripts/xg_dispute_analysis.py --sage-picks data/money_benchmark/scores/sage_truncated3.picks.jsonl

The first parses XG's rollouts into data/money_benchmark/xg_results/rollout.jsonl; the second prints the disputed-position report — every disagreement scored against the Sage and XG rollouts in turn — from Sage 3T's recorded picks.

Match play is not yet covered by this method: it needs an XG full rollout of the match positions, which we have not run. Match strength is covered instead by the Rollout PR study (Section 3) and the real-match agreement study (Section 5).

Head-to-head games

Sage and XG can also play each other directly, with XG's moves taken from its own analysis. Sage plays both sides of an unlimited game; the transcript is batch-analysed in XG; the first opponent decision that differs from XG's top choice is replaced by XG's move and the game replayed from there with fresh dice, until every opponent move is the one XG would have made. The runner plays a series of such games (seeds continue from logs/sage_vs_xg.txt) and prints the mean points per game with its standard error:

python scripts/run_sage_vs_xg_games.py 100 --level 3P   # Sage 3P vs XG; levels 1P-4P, 1T-3T
python scripts/play_sage_vs_xg.py 3P --seed 7             # one game, iterations under logs/play_sage_vs_xg/

XG must be running on the Windows desktop: every iteration drives its Batch Analyze dialog through scripts/xg_batch_analyze.py, an Anthropic Computer Use agent (set ANTHROPIC_API_KEY), which takes the mouse and keyboard for about a minute per iteration. A game converges in a handful of iterations; XG's "too good to double" verdict counts as no double.

Backgames, containment games and the snake

The thirteen family benchmarks live under backgame_ref_positions/benchmark/ — one <family> starting.txt of seed positions and one <family> rollout.jsonl reference each — and are scored by scripts/score_backgame_pr.py, which reports the PR beside the share of picks that are rollout-graded.

python scripts/score_backgame_pr.py --category "21 backgame" --level truncated3
python scripts/score_backgame_pr.py --category containment --level 3ply
# XG: export every decision as a one-decision native .xg game, then per level (XG running): analyse, wait, score
python scripts/export_folder_benchmark_xg.py generate
python scripts/xg_folder_batch.py stamp
python scripts/xg_folder_batch.py analyze --level xg3ply   # drives XG's Batch Analyze dialog (pywinauto)
python scripts/xg_folder_batch.py wait --level xg3ply
python scripts/xg_folder_batch.py score --level xg3ply
Open Sage is licensed under MPL-2.0. The shipped benchmarks are the exact references behind the Section 3 tables; both the XG-reference comparison and the disputed-position results regenerate from them once XG's Batch Rollout has produced the matching .xg files.