Declared vs Measured RTP: What the Lab’s Tests Show
Every slot prints an RTP on its help screen. This report puts that figure next to what ten thousand logged rounds per game actually returned, for every test in the lab, and explains why the two rarely match and what the gap can and cannot tell you.
Two numbers that are not the same thing
The declared RTP is a design value: the share of all stakes a game is built to return over a volume of play that no single session reaches. The measured RTP is arithmetic on one run: total win divided by total bet over 10,000 rounds at a fixed stake on the official demo server. Across the lab that comparison now covers 68 tests and 680,000 logged rounds. The average declared figure of the tested builds is 96.39%, the average measured return 95.8%.
The gap between the two averages is small; the gap inside any single test is not. The lowest measured return in the set is 77.1%, the highest 144.9%, on games whose declared figures sit within a couple of points of each other. That spread is the subject of this report.
The picture across the lab
Each bar is one test, sorted by measured return. A bar runs from the figure printed in the game’s help screen to the figure the run actually returned: teal where the run finished above the printed value, amber where it finished below, with the dashed line marking the average declared RTP of the set. Hovering a bar shows the approximate 95% interval of that run: the range in which the true design value would be expected to sit, given the size and variance of the sample. 42 of 68 tests came in below their declared value and 26 came in above. That is close to an even split, which is what a fair set of finite runs looks like; a lab that only ever measured below the printed figure would be measuring something other than variance.
The histogram shows the same runs by measured value alone. The bulk sits in the bins around the declared average, with a tail on each side made of a handful of games where a single feature round worth several hundred bets, or its absence, moved the run by several points. Those tails are not the loosest or tightest games; they are the most volatile ones.
Which games landed furthest from their figure
| game | declared | measured | gap, pts | 95% band | declared inside |
|---|---|---|---|---|---|
| Wild Beach Party | 96.53% | 144.9% | +48.3 | 71.3–218.4% | yes |
| Book of Golden Sands | 96.46% | 120.6% | +24.2 | 85.4–155.8% | yes |
| Bonanza Trillion | 97.17% | 120.8% | +23.6 | 88.3–153.3% | yes |
| Buffalo King Megaways | 96.52% | 116.5% | +20.0 | 87.2–145.8% | yes |
| Gates of Olympus 1000 | 96.50% | 115.9% | +19.4 | 86.5–145.4% | yes |
| Starlight Princess 1000 | 96.50% | 111.3% | +14.8 | 59.3–163.4% | yes |
| Twilight Princess | 96.08% | 110.9% | +14.8 | 84.8–136.9% | yes |
| The Hand of Midas | 96.54% | 108.9% | +12.4 | 83.9–134.0% | yes |
| game | declared | measured | gap, pts | 95% band | declared inside |
|---|---|---|---|---|---|
| Big Bass – Hold & Spinner | 96.07% | 77.1% | -18.9 | 57.7–96.6% | yes |
| Forge of Olympus | 96.25% | 79.7% | -16.6 | 69.4–90.0% | no |
| Gates of Olympus | 96.50% | 80.2% | -16.3 | 68.6–91.8% | no |
| Big Bass Bonanza | 96.71% | 81.2% | -15.5 | 66.8–95.6% | no |
| Fire Strike 2 | 96.53% | 82.6% | -13.9 | 71.8–93.5% | no |
| Big Bass Splash | 96.71% | 83.1% | -13.6 | 66.3–99.9% | yes |
| Big Bass Day at the Races | 96.07% | 82.7% | -13.4 | 59.2–106.2% | yes |
| The Dog House | 96.51% | 83.6% | -12.9 | 67.3–99.9% | yes |
A gap of ten points or more in either direction almost always has a round number attached: one natural feature that paid a four-figure multiple, or a stretch of a thousand rounds where the feature never came. The review of each game states which it was. What the table does not show is a game that measures far from its declared figure across repeated runs. That would be a different finding, and it is what re-tests are for.
Why 10,000 spins leaves a wide band
The interval next to every measured RTP is the honest part of the number. For a low-volatility game, where most rounds pay a small multiple, ten thousand rounds pin the mean to within a few points. For a high-volatility game, where the design return is carried by rare rounds of hundreds of bets, the same sample leaves a band of thirty points or more, because whether one or two of those rounds happened inside the window decides the result. In 63 of 68 tests the declared figure sits inside the measured band, which is the expected outcome for a correctly declared game; the 5 tests where it does not are the ones at the two ends of the histogram.
This is also why a single measured RTP is never presented here as evidence that a game is set higher or lower than it says. One run cannot establish that. What it can establish is what the run returned, and how far from the printed figure a real session of that length can land.
The declared figure is a version, not a constant
Many tested games ship in several RTP versions and let the operator choose one. The declared column in the lab records the profile the demo server reported for the build it played, which is normally the headline version. The copy at a given casino may be a lower one, identical in every visible way. This report therefore describes the headline builds; a casino running a 94% version of the same game should be expected to measure correspondingly lower over a comparable run, and the only way to know which version is in front of you is the game’s own information screen at that casino.
By provider
| provider | tests | mean measured RTP | mean declared | mean hit rate | mean buy value |
|---|---|---|---|---|---|
| Pragmatic Play | 60 | 95.6% | 96.39% | 26.2% | 0.93x |
| Relax Gaming | 7 | 94.1% | 96.21% | 21.7% | 0.81x |
| BGaming | 1 | 120.8% | 97.17% | 31.1% | 1.06x |
The provider grouping is here because the question is asked often, not because the sample supports a verdict: the lab is weighted towards one studio, and the measured averages by provider move with every batch of tests. The declared averages are the steadier column, and they say more about each studio’s default RTP profile than the measured ones say about anything.
What this report will look like as tests are added
Every figure above is recalculated from the current test set when the page loads, and the note under each block names that set. Two things are expected as the lab grows: the measured average should stay close to the declared average, and the share of tests with the declared figure inside their band should stay near the level a 95% interval implies. If either drifts, that drift is the next report.
Data behind this report
68 tests · 2026-09-03 to 2026-09-16Every table and chart on this page is computed from the 68 slot tests published on Slotester at the time you opened it: 680,000 logged rounds at a fixed bet, 2,900 bonus purchases in separate sessions, 2,705 natural bonus triggers. A test enters the set the moment its review is published, so the figures move as the lab grows; the badge at the top names the test set used for this view.
The written analysis describes the evidence as it stood on the publication date and is revised when a new batch of tests changes the picture. Per-test data is not summarised away: each review in the catalog links its full spin log with a SHA-256 checksum, and the procedure behind every run is on the how we test page.
A measured figure is what one 10,000-round run returned, not the game's long-run design value. Intervals are approximate 95% bands for the mean of that run. Averages across games weight every test equally; medians are given where a few large results would otherwise carry the mean.