COPA / EVALUATION
Model Evaluation — LLM vs Caco Model
QUARTER-FINAL — LLM self-evaluation (86 starters)
MAE — mean / median34.7k / 15.2k
CRPS25.4k
Rank — Spearman ρ0.71 (R16 was 0.42)
NDCG@400.74
Direction accuracy77%
80% interval coverage88%
Best round of the tournament. With only eight teams and clear favourites, the QF was
the model's cleanest read: rank skill ρ=0.71, direction 77%, and interval
coverage 88%. A clean rebound from the upset-driven R16 — the field reverted to form and the
odds-derived projections tracked it closely. (Semi-final skipped — no predictions were generated that round.)
ROUND OF 16 — LLM self-evaluation (177 starters)
MAE — mean / median51.5k / 28.8k
CRPS36.6k
Rank — Spearman ρ0.42 (R32 was 0.62)
NDCG@400.51
Direction accuracy63%
80% interval coverage73% (R32 was 72%)
The upset round. Brazil lost to Norway, Belgium put four past the USA, and Ounahi (405k)
and De Ketelaere (447k) hauled from midfield — outcomes no odds-derived model foresaw. Rank skill fell to
ρ=0.42 (from 0.62) and R² to 0.11: the ordering was genuinely hard. What held
was the distribution — 73% interval coverage, in line with R32 — so the wide predictive
bands still contained the chaos even when the point ranking missed.
ROUND OF 32 — LLM self-evaluation (411 starters)
MAE — mean / median40.6k / 20.5k
CRPS29.1k
Rank — Spearman ρ0.62
NDCG@400.64
Direction accuracy63%
80% interval coverage72%
The knockout upgrades landed: start-rate weighting (R1/R2 over dead-rubber R3) and
difficulty-adjusted goal shares gave the model its cleanest read of the tournament — rank skill
ρ=0.62, coverage 72%. (R32 & R16 are scored on the same basis: projected
starters, p_play ≥ 0.5.)
ROUND 3 — LLM self-evaluation (732 players)
MAE — mean / median44.6k / 45.2k
CRPS34.1k
Rank — Spearman ρ0.36
80% interval coverage59% (R1 43% → R2 54% → R3)
R3 was the knockout-adjacent round: heavy dead-rubber rotation and freak hauls (Dembélé 602k) cut
rank skill to ρ=0.36 and the median didn't beat the mean (the rotation dampener pushed p50 too low for
stars who ended up playing, e.g. Messi). Interval coverage kept climbing to 59% — widening the
predictive distribution remains the top model fix.
ROUND 2 — LLM self-evaluation (742 players)
MAE — median (p50) point44.5k
MAE — mean point45.0k
CRPS34.2k
Rank — Spearman ρ0.44
NDCG@400.56
80% interval coverage54% (R1 was 43%)
R2 validated the upgrades: the median point beat the mean on MAE
(44.5k vs 45.0k, the Gneiting loss-matching fix); interval coverage rose
43% → 54% (still short of 80% — widening the distribution is the next fix);
rank skill held at ρ=0.44. Caco has not published Round-2 predictions yet, so R2 is a self-evaluation of the LLM only. Head-to-head resumes when Caco posts R2.
Round 1 — head-to-head vs Caco Model
PLAYERS SCORED COMMITTED 2026-06-11
Both models scored against the same live actuals (players whose match has been played).
Caco Model = the public
Copa Analytics dashboard ("Ours" column there);
Large Lindmarker Model (LLM) = ours.
Metric set follows the forecast-evaluation literature: proper scoring (
CRPS),
loss-matched point error (
Gneiting 2011), an
asinh signed-log error (proportional in the tails, punishes sign-flips near 0),
skill vs a position-mean baseline (
MASE-style),
directional accuracy + the
Pesaran–Timmermann test, and rank correlation / NDCG.
★ Best measure — CRPS
Large Lindmarker Model (LLM)—
Caco Modelnot yet comparable
CRPS is the one to watch: it scores the whole predictive distribution, rewards being both
calibrated and sharp, and is the fairest measure given how much randomness there is in a single round.
We can compute it for LLM because the model outputs a full distribution (p10/p50/p90 per player).
For a head-to-head, Caco Model needs to publish a predictive distribution — e.g. p10 / p50 / p90
(or a standard deviation) for each player, not just one point number. Until then the table below uses
point- and rank-based metrics, which are all a point forecast supports.
Other metrics (point & rank — work on point forecasts)
Per-player — actual vs each model (sortable)
| Player | Team | Pos |
Actual | LLM | Caco |