The calculator takes both team previews and makes two kinds of guesses. It predicts what your opponent will bring and lead. It also ranks your own leads and fours by how similar choices have done on the ladder. This note checks both against games the calculator had never seen.
How it was tested
A fair test has to hide the answers. So the test was split by date:
- The calculator’s tables were rebuilt using only games up to Sep 24, 206,526 games in all.
- It was then scored on every game from Sep 25 to Sep 28: 71,572 games, or 143,144 sides. The last day was partial.
- Its handful of tuning weights were chosen earlier still, on Sep 23–24 games, using tables built only from games before those.
None of the test games touched the tables or the weights. A model scored on its own training data always looks better than it is.
Guessing their lead
There are 15 possible lead pairs from a team of six. A blind guess is right 1 time in 15, or 6.7%. A smarter baseline picks whichever of their possible pairs has led most often on the ladder. The calculator does better than both.
| Prediction | Calculator | Baseline |
|---|---|---|
| Their lead pair, top guess | 22.3%CI 22.1–22.5 · n 143k | 12.9%Most common pairCI 12.8–13.1 · n 143k |
| Their lead pair, in top three | 47.3%CI 47.1–47.6 · n 143k | 32.2%Most common pairCI 32.0–32.5 · n 143k |
| Their brought mons in predicted four | 74.6%CI 74.3–74.8 · n 143k | 70.4%Top four by bring rateCI 70.2–70.7 · n 143k |
| Their exact four (full-reveal sides) | 17.8%CI 17.6–18.1 · n 90k | 9.3%Top four by bring rateCI 9.1–9.5 · n 90k |
Its single top guess was right 22.3% (95% CI 22.1–22.5, n 143,144), against 12.9% (95% CI 12.8–13.1, n 143,144) for the baseline. The real lead was in its top three 47.3% (95% CI 47.1–47.6, n 143,144), against 32.2% (95% CI 32.0–32.5, n 143,144).
Put plainly: it is wrong about the exact lead most of the time. Leads depend on the matchup and the player, and a species-level model can’t see sets, items or habits. What it does well is narrow 15 options to a short list.
Guessing their four
Of the Pokémon each opponent actually showed, 74.6% (95% CI 74.3–74.8, n 143,144) were in the calculator’s predicted four. Picking the four species with the highest bring rates gets 70.4% (95% CI 70.2–70.7, n 143,144). The gap is smaller, since bring rates alone say a lot.
The exact-four number is stricter. In games where the opponent showed all four, the calculator named the whole four 17.8% (95% CI 17.6–18.1, n 89,686), about twice the baseline’s 9.3% (95% CI 9.1–9.5, n 89,686).
Recall slightly understates itself: a Pokémon brought but never sent out counts as “not brought”. Full-reveal games avoid that and give almost the same recall.
Are the win-rate estimates honest?
The calculator also puts a number on each of your options, such as “this lead has won about 54%”. A number like that is only useful if it means what it says. So the test sides were grouped by the calculator’s estimate and compared with what they actually won.
- Est. 35.136.7
- Est. 37.137.6
- Est. 39.139.7
- Est. 41.141.1
- Est. 43.043.4
- Est. 45.046.4
- Est. 47.047.9
- Est. 49.050.1
- Est. 51.052.1
- Est. 53.055.3
- Est. 55.057.3
- Est. 56.959.3
- Est. 58.960.1
- Est. 60.963.2
- Est. 62.963.9
The estimates track reality closely. In every 2-point bin with at least 1,000 test sides, the observed win rate is within about 3 points of the estimate. Above 50% it is usually 1 to 2 points higher than the estimate. So when the calculator is wrong, it is usually too cautious, not too optimistic. That is by design: thin data is pulled toward the average, which trims extreme estimates.
Do its picks win?
This is the question the data can answer least well.
- Led the top pick55.6
- Led a top-3 pick54.3
- Led anything else49.3
- Led a pair ranked 9–1546.5
Sides that happened to lead the calculator’s top pick won 55.6% (95% CI 54.8–56.4, n 15,275). Every other lead won 49.3% (95% CI 49.1–49.6, n 127,557). Sides whose lead was ranked 9th to 15th won 46.5% (95% CI 46.1–46.9, n 54,534). For fours, sides that brought the top pick won 54.8% (95% CI 53.9–55.8, n 9,977), against 49.4% (95% CI 49.0–49.7, n 79,709) for any other four.
That looks like a 6-point edge. It is not evidence that following the calculator adds 6 points. The calculator was not public when these games were played. These are players who chose, on their own, what the calculator would later have picked. Players who choose what the data favours are probably better players to begin with. The estimates are also partly built from the win rates of similar players in earlier games. So this check shows that the ranking agrees with what wins. It does not show that switching to it causes wins. A real test would need players randomly assigned to follow it, which a ladder dataset can’t provide.
What it can’t do
- Species only. It knows nothing about items, abilities, moves or spreads. Two teams with the same six can play completely differently.
- Pairs only. It adds up pairwise effects (this species next to that one, this one against that one) as if they were independent. Three-way interactions are not modelled.
- Leads are chosen after preview. A lead’s win rate comes from players who chose it into the teams they saw, so a lead picked only into good matchups will look better than it is.
- Player clustering. A niche pairing carried by a few players looks more certain than it is.
- All Elo only. The ≥ 1300 sample is too thin for cross-team cells.
- A moving target. The test covers four days of a new regulation. Accuracy may drop as the meta shifts.
For scale, the mimikyu model for the previous regulation, M-B, reported 23.3% lead top-1 and 66.1% bring recall, on different data and a different split. Compare loosely.
The fair summary: a useful short list and an honest estimate, not an oracle. Use it to narrow your options, then play the game in front of you.
Every number here comes from the snapshot above, frozen in this file. Win rates carry 95% Wilson intervals, which assume independent games; see player clustering.