The full forecast record
Every published forecast is scored once its match is played: no filtering by league, confidence or outcome. The rules were registered before the first result was known (protocol v1). Lower Brier score and log loss are better.
Accuracy
| Market | Matches | Brier | Log loss | First forecast Brier | First forecast log loss |
|---|---|---|---|---|---|
| 1X2 | 0 | – | – | – | – |
| Over/under 2.5 | 0 | – | – | – | – |
| Both teams to score | 0 | – | – | – | – |
Model vs market
Paired on the same matches. Difference is model minus benchmark: negative means the model scored better. No conclusion is drawn while the 95% interval includes zero.
| Benchmark | Market | Metric | Matches | Model | Benchmark | Difference | 95% interval | Conclusion |
|---|---|---|---|---|---|---|---|---|
| Market at forecast time | 1X2 | Brier | 0 | – | – | – | – | insufficient data |
| Market at forecast time | 1X2 | Log loss | 0 | – | – | – | – | insufficient data |
| Market at forecast time | Over/under 2.5 | Brier | 0 | – | – | – | – | insufficient data |
| Market at forecast time | Over/under 2.5 | Log loss | 0 | – | – | – | – | insufficient data |
| Market at forecast time | Both teams to score | Brier | 0 | – | – | – | – | insufficient data |
| Market at forecast time | Both teams to score | Log loss | 0 | – | – | – | – | insufficient data |
| Closing market (market average) | 1X2 | Brier | 0 | – | – | – | – | insufficient data |
| Closing market (market average) | 1X2 | Log loss | 0 | – | – | – | – | insufficient data |
| Closing market (market average) | Over/under 2.5 | Brier | 0 | – | – | – | – | insufficient data |
| Closing market (market average) | Over/under 2.5 | Log loss | 0 | – | – | – | – | insufficient data |
| Closing market (market average) | Both teams to score | Brier | 0 | – | – | – | – | insufficient data |
| Closing market (market average) | Both teams to score | Log loss | 0 | – | – | – | – | insufficient data |
Calibration
When the model says 60%, does it happen about 60% of the time? Points on the dashed line are perfectly calibrated; larger points hold more predictions.
1X2
0 predictionsPredicted % (horizontal) vs observed % (vertical).
Over/under 2.5
0 predictionsPredicted % (horizontal) vs observed % (vertical).
Both teams to score
0 predictionsPredicted % (horizontal) vs observed % (vertical).
Market movement
Between the forecast and kickoff, did the market move towards the model? This is a closing-line comparison without betting.
| Selection | Matches | Mean movement towards model | Share of matches moving towards model |
|---|---|---|---|
| Home win | 0 | – | – |
| Over 2.5 | 0 | – | – |
Newcomers
Matches with a team that did not play in the league last season (usually newly promoted) are reported separately, as registered before the first result, and always next to the full population.
| Matches | Count | 1X2 Brier | 1X2 log loss | Over 2.5 Brier | 1X2 log loss vs market at forecast |
|---|---|---|---|---|---|
| All matches | 0 | – | – | – | – |
| With a newcomer | 0 | – | – | – | – |
| Without a newcomer | 0 | – | – | – | – |
By league
Every evaluated match
Research selections Research, not advice
oddsxi-goal-model p6-v5A research record of pre-match disagreements between the model and the market. Whenever the model probability was at least 3 percentage points above the market consensus (margin removed), the case was stored before kickoff and can never be changed. After kickoff it appears here with the result. The question is simple: do these cases come true more often than the market expected? Rules registered before the first result; no conclusion while the interval includes zero.
| Model above market by | Cases | Correct–wrong(–void line) | Came true | 95% interval | Market expected | Above market | 95% interval |
|---|---|---|---|---|---|---|---|
| 3+ points | 0 | 0–0 | – | – | – | – | – |
| 5+ points | 0 | 0–0 | – | – | – | – | – |
| 8+ points | 0 | 0–0 | – | – | – | – | – |
Baseline: oddsxi-goal-model p6-v1
0 evaluated · 86 pendingThe previous version, which treated newly promoted teams as league-average teams, is still published and scored on the same matches, so the newcomer fix is also tested on live data.
| Market | Matches | Brier | Log loss | First forecast Brier | First forecast log loss |
|---|---|---|---|---|---|
| 1X2 | 0 | – | – | – | – |
| Over/under 2.5 | 0 | – | – | – | – |
| Both teams to score | 0 | – | – | – | – |
| Matches | Count | 1X2 Brier | 1X2 log loss | Over 2.5 Brier | 1X2 log loss vs market at forecast |
|---|---|---|---|---|---|
| All matches | 0 | – | – | – | – |
| With a newcomer | 0 | – | – | – | – |
| Without a newcomer | 0 | – | – | – | – |
Cards model Beta
0 evaluated · 86 pendingoddsxi-cards-model p7-v1 against oddsxi-cards-baseline p7-v1, a league-average baseline registered before the first forecast.
| Measure | Model | Baseline | Difference | 95% interval |
|---|---|---|---|---|
| Log loss, full distribution | – | – | – | – |
| Brier, over 3.5 cards | – | – | – | – |
| Brier, over 4.5 cards | – | – | – | – |