COPA / EVALUATION

Model Evaluation — LLM vs Caco Model

PredictionsStandingsEvaluationCaco-gate
QUARTER-FINAL — LLM self-evaluation (86 starters)
MAE — mean / median34.7k / 15.2k
CRPS25.4k
Rank — Spearman ρ0.71 (R16 was 0.42)
NDCG@400.74
Direction accuracy77%
80% interval coverage88%
Best round of the tournament. With only eight teams and clear favourites, the QF was the model's cleanest read: rank skill ρ=0.71, direction 77%, and interval coverage 88%. A clean rebound from the upset-driven R16 — the field reverted to form and the odds-derived projections tracked it closely. (Semi-final skipped — no predictions were generated that round.)
ROUND OF 16 — LLM self-evaluation (177 starters)
MAE — mean / median51.5k / 28.8k
CRPS36.6k
Rank — Spearman ρ0.42 (R32 was 0.62)
NDCG@400.51
Direction accuracy63%
80% interval coverage73% (R32 was 72%)
The upset round. Brazil lost to Norway, Belgium put four past the USA, and Ounahi (405k) and De Ketelaere (447k) hauled from midfield — outcomes no odds-derived model foresaw. Rank skill fell to ρ=0.42 (from 0.62) and R² to 0.11: the ordering was genuinely hard. What held was the distribution — 73% interval coverage, in line with R32 — so the wide predictive bands still contained the chaos even when the point ranking missed.
ROUND OF 32 — LLM self-evaluation (411 starters)
MAE — mean / median40.6k / 20.5k
CRPS29.1k
Rank — Spearman ρ0.62
NDCG@400.64
Direction accuracy63%
80% interval coverage72%
The knockout upgrades landed: start-rate weighting (R1/R2 over dead-rubber R3) and difficulty-adjusted goal shares gave the model its cleanest read of the tournament — rank skill ρ=0.62, coverage 72%. (R32 & R16 are scored on the same basis: projected starters, p_play ≥ 0.5.)
ROUND 3 — LLM self-evaluation (732 players)
MAE — mean / median44.6k / 45.2k
CRPS34.1k
Rank — Spearman ρ0.36
80% interval coverage59% (R1 43% → R2 54% → R3)
R3 was the knockout-adjacent round: heavy dead-rubber rotation and freak hauls (Dembélé 602k) cut rank skill to ρ=0.36 and the median didn't beat the mean (the rotation dampener pushed p50 too low for stars who ended up playing, e.g. Messi). Interval coverage kept climbing to 59% — widening the predictive distribution remains the top model fix.
ROUND 2 — LLM self-evaluation (742 players)
MAE — median (p50) point44.5k
MAE — mean point45.0k
CRPS34.2k
Rank — Spearman ρ0.44
NDCG@400.56
80% interval coverage54% (R1 was 43%)
R2 validated the upgrades: the median point beat the mean on MAE (44.5k vs 45.0k, the Gneiting loss-matching fix); interval coverage rose 43% → 54% (still short of 80% — widening the distribution is the next fix); rank skill held at ρ=0.44. Caco has not published Round-2 predictions yet, so R2 is a self-evaluation of the LLM only. Head-to-head resumes when Caco posts R2.

Round 1 — head-to-head vs Caco Model

PLAYERS SCORED COMMITTED 2026-06-11
Both models scored against the same live actuals (players whose match has been played). Caco Model = the public Copa Analytics dashboard ("Ours" column there); Large Lindmarker Model (LLM) = ours. Metric set follows the forecast-evaluation literature: proper scoring (CRPS), loss-matched point error (Gneiting 2011), an asinh signed-log error (proportional in the tails, punishes sign-flips near 0), skill vs a position-mean baseline (MASE-style), directional accuracy + the Pesaran–Timmermann test, and rank correlation / NDCG.
★ Best measure — CRPS
Large Lindmarker Model (LLM)
Caco Modelnot yet comparable
CRPS is the one to watch: it scores the whole predictive distribution, rewards being both calibrated and sharp, and is the fairest measure given how much randomness there is in a single round. We can compute it for LLM because the model outputs a full distribution (p10/p50/p90 per player). For a head-to-head, Caco Model needs to publish a predictive distribution — e.g. p10 / p50 / p90 (or a standard deviation) for each player, not just one point number. Until then the table below uses point- and rank-based metrics, which are all a point forecast supports.

Other metrics (point & rank — work on point forecasts)

MetricLLMCaco ModelEdge

Per-player — actual vs each model (sortable)

PlayerTeamPos ActualLLMCaco