betaider.com

Track record

Tennis: what the model predicted, against what actually happened

What you are looking at
Two records, answering different questions. Historical replays every prediction the walk-forward backtest made — the model only ever saw matches played before the one it was predicting, so these are genuine forecasts, and there are enough of them to mean something. Live is picks this app actually published, settled after the fact. It starts empty and grows.

Historical uses the raw model probability rather than the calibrated one, because calibration is itself derived from this history — scoring against it would be marking its own homework.

Historical — out of sample

Selections settled
380,696
Model said
51.9%
Actually won
51.6%
Calibration gap
-0.36pp

A gap near zero means the model's probabilities mean what they say: things it calls 70% happen about 70% of the time. That is the single most important number on this site.

Calibration by confidence

Model saidSelectionsPredicted ActualGapPredicted vs actual
0.05–0.10 25,336 7.6% 11.9% +4.32
0.10–0.15 22,710 12.8% 17.2% +4.44
0.15–0.20 18,428 17.6% 23.3% +5.73
0.20–0.25 15,842 22.5% 26.9% +4.38
0.25–0.30 20,314 27.6% 28.4% +0.84
0.30–0.35 17,608 32.6% 36.2% +3.58
0.35–0.40 18,375 37.6% 39.6% +1.99
0.40–0.45 18,855 42.7% 42.4% -0.33
0.45–0.50 32,458 48.2% 46.0% -2.16
0.50–0.55 32,067 51.8% 52.9% +1.07
0.55–0.60 17,785 57.2% 56.2% -1.04
0.60–0.65 16,698 62.3% 59.5% -2.81
0.65–0.70 15,218 67.4% 63.8% -3.60
0.70–0.75 13,158 72.5% 68.8% -3.75
0.75–0.80 10,796 77.6% 71.0% -6.53
0.80–0.85 14,129 82.4% 76.2% -6.21
0.85–0.90 13,069 87.4% 80.4% -7.05
0.90–0.95 21,096 92.5% 87.9% -4.63
0.95–1.00 36,754 98.4% 96.3% -2.16
predicted actual

By market

Market familySelectionsPredicted ActualGap
Games handicap 137,637 59.2% 58.7% -0.44
Total games 117,491 50.4% 50.3% -0.03
Set handicap 26,984 67.5% 66.9% -0.56
Straight sets 17,876 39.3% 40.2% +0.82
Set score 17,721 25.6% 23.7% -1.89
First set / match 17,532 25.9% 25.8% -0.10
Deciding set 9,108 50.0% 48.0% -2.04
Odd/Even 9,108 50.0% 50.0% +0.00
Total sets 9,108 50.0% 48.0% -2.04
First set 9,107 50.0% 50.0% +0.00
Match winner 9,024 50.4% 50.4% -0.02

A family with a large negative gap is one the model oversells. The radar's calibration already corrects for this before ranking, but it is worth knowing which markets to trust least.

By league

CompetitionSelectionsPredicted ActualGap
WTA results-only 238,147 51.7% 51.7% -0.05
ATP results-only 142,549 52.3% 51.4% -0.87

Live record

Model
ai-blend-live
Settled
13,002
Model said
36.4%
Actually won
36.6%
Model
dc-xg-live
Settled
13,138
Model said
36.3%
Actually won
36.5%
Model
deep-live
Settled
1,786
Model said
36.7%
Actually won
36.7%

Profit at fair odds

Flat 1 unit per selection
+27476
Per selection
+0.0722

This is not a profit claim. Every bet is settled at the model's own fair odds, where a perfectly calibrated model scores exactly zero by construction. It is a second calibration measure, and it is more sensitive than the headline gap because a longshot priced at 7% pays about 13× — so small errors on unlikely outcomes cost far more than the same error on a favourite. A negative figure alongside a near-zero overall gap points to the familiar favourite–longshot pattern: the model slightly oversells long odds and slightly undersells short ones. Real bookmaker prices are shorter than fair odds, so actual betting would do worse than this.

betaider.com

Sport

Tennis sections

Preferences

Appearance
Aider Reading the Tennis model · online
Open the Radar
Bet slip
Open full page
to move to open esc to close See all results

On this page

Track record

What this page is for

The scoreboard. Not what the model thinks it can do, but what happened when it was asked in advance and the matches were then played.

How to read it

Everything here is out of sample: each prediction was made before the match it describes. Break the record down by market and league before drawing a conclusion from the headline — a model can be genuinely good at one league and useless at another, and an average over both hides it.

Things to watch

Sample size is the trap. A market with thirty settled picks tells you almost nothing; a run of good results over a few weeks is well within what luck produces. Football is low-scoring enough that the better side loses often, and a fair chunk of every result is simply not predictable.

Where to go next

Profit at fair odds, at the foot of the page, is the harshest test on the site: it asks whether the edge survives being paid for.

Press esc to close Read the full guide