# 2025 Model Comparison & Market Anomaly Hunt

**Generated:** 2026-08-05 | **Season:** 2025 | **Train:** 2016–2024 (4,520 games) | **Test:** 2025 (550 games)

**Model-comparison correction:** 2026-09-08. The anomaly-hunt section below
retains its original results; it uses a separate model-edge strategy.

## Model comparison: probability scores and spread performance

Every model trained on 2016-2024 and predicted 2025. These seasons informed
feature development, so this is retrospective evaluation, not an untouched
holdout. Coin-flip baselines are Brier 0.25 and log loss 0.693.

| Model | Brier | Log loss | Winner accuracy | ATS cover rate |
|---|---|---|---|---|
| 05 LogisticRegression | **0.1830** | **0.5401** | 0.7236 | 0.5184 |
| 07 Stacked ensemble | 0.1850 | 0.5489 | **0.7418** | **0.5294** |
| 02 RF team points | 0.2140 | 0.7083 | 0.7164 | 0.5129 |
| 03 XGBoost win prob | 0.2245 | 0.6897 | 0.6800 | 0.4945 |
| 01 Ridge margin | 0.3054 | 0.8906 | 0.5782 | 0.5110 |
| Market favorite | | | 0.7345 | 0.5110 |

Winner accuracy covers 550 games. ATS backs each predicted winner at the
recorded spread and grades 544 games, excluding six pushes. The market ATS
baseline backs the spread favorite and also excludes pick'em lines.
This strategy is not the dashboard's model-edge strategy.

The old `ats` and `market_ats` fields incorrectly counted straight-up wins:
the ensemble's 74.18% was winner accuracy, not ATS performance. The scorer
now uses final score plus spread and excludes pushes from its denominator.
`tests/test_model_comparison.py` covers winners that fail to cover and the
push denominator. `exports/model_comparison_2025.json` was regenerated by
`analysis/model_comparison.py`; probability scores and winner accuracy did
not change. The separate dashboard `betting.ats` implementation was not changed.

The ensemble has the highest winner accuracy among these five models;
logistic regression has the lowest Brier and log loss. The ensemble's
52.94% ATS point estimate is only slightly above 52.38% break-even at -110
and does not establish a profitable edge. Brier and log loss measure overall
probability quality, not calibration alone. fastai was not evaluated; SHAP
is an interpretability method, not a competing predictor.

## Market anomaly hunt — did the documented biases replicate in 2025?

The market-efficiency literature reports persistent structural biases: favorites overpriced (Sinkey & Logan 2009, 11k games 1985–2003), home underdogs 7+ profitable (Paul et al. 2003, 25 years), big favorites exploitable. Tested against 567 lined 2025 games. Break-even at -110 = 52.38%.

| Test | N | Cover rate | Binom p | Verdict |
|---|---|---|---|---|
| All games | 567 | 50.6% | 0.80 | Efficient |
| Favorites | 339 | 51.0% | 0.74 | No bias |
| Underdogs | 228 | 50.0% | 1.00 | No bias |
| Favorites 14+ | 92 | 50.0% | 1.00 | No bias |
| Favorites 21+ | 42 | 42.9% | 0.44 | Underdogs cover, n too small |
| Home underdogs 7+ | 114 | 50.9% | 0.93 | No bias |
| \|edge\| 0–3 (model pick) | 232 | 52.0% | 0.60 | No edge |
| \|edge\| 3–7 (model pick) | 173 | 51.2% | 0.82 | No edge |
| \|edge\| 7–14 (model pick) | 102 | 43.6% | 0.23 | Model loses here |
| \|edge\| 14+ (model pick) | 9 | 77.8% | 0.18 | Only winner — n=9, noise |

**Read:** The classic anomalies did **not** replicate in 2025 — the market was efficient on every documented test. The one signal that survives is the model's own 14+ edge bucket (77.8%), but at n=9 it is a whisper, not a track record. The 7–14 bucket (43.6%) remains the model's known weak spot against the number. Team-totals censoring bias (Arscott 2023) untested — no totals data in cfb.db.

## What this means

1. The ensemble's 74.2% winner accuracy and 52.9% winner-backed ATS rate measure different outcomes. Neither establishes a market-beating model.
2. Logistic regression has better Brier and log loss on this retrospective 2025 comparison; probability reliability needs a separate calibration check.
3. The anomaly-hunt section uses the separate EPA model-edge strategy. Its nine-game 14+ bucket is not evidence of a reliable betting edge.
4. Compare proposed model changes on the same chronological game cohorts before choosing a model.

*Model output, not betting advice. Historical hit rates are small-sample backtests; several edge buckets lose money at -110. Flat units only.*
