How to read a leaderboard

A ranked table invites a quick conclusion. This page is about the places where the quick conclusion is wrong.

The first half holds for every board. The second half takes each live board in turn — Transfers and Team Value — and what its own columns mean. The terms used here are defined in the glossary.

Zero is the interesting number

Not one, and not the top of the table.

Every skill score is measured against a baseline: the answer anyone could give without a model of their own. Zero means a participant did no better than that — which is a real thing to know, freely available, and requires no model at all. Above zero is the whole claim.

  • On Transfers, zero is the league base rate for the player's role. Above it, a participant is extracting something about individual players that the league average does not contain.
  • On Team Value, zero is the scoreline. Above it, a participant is seeing something in matches that goal difference does not.

That also makes the scale unintuitive. Skill scores are not percentages and not accuracy, and the two boards' scales are not comparable with each other: +0.21 is a strong Transfers result, while a season of Team Value resolves differences of about five hundredths.

Negative is shown as negative. It means worse than the baseline, which is a real outcome for a confident model pointed at the wrong questions, and clamping it to zero would hide the difference between a bad model and a cautious one.

The four misreadings

1. Reading the ranking before the interval

Early in a season the interval is the story and the point estimate is barely a story at all. A handful of questions with little football behind them cannot separate a good model from a lucky one, and the interval says so.

If two intervals overlap heavily, the board is telling you it cannot yet order those two rows — even though it has to print them in some order.

The Team Value board says this in words. Every row reads ahead of, level with or behind the scoreline, depending on whether its interval clears zero, and two rows that both read level have not been separated from the scoreline, whatever their rank numbers say. Read the words before the number.

2. Reading overlap as "no difference"

The reverse trap, and the subtler one. Two overlapping intervals do not mean two participants are indistinguishable.

Participants answer identical questions, so their errors are correlated, and the difference between two scores is far better determined than either score. On a real twelve-case Transfers cohort, two participants with heavily overlapping intervals — 0.498 [0.387, 0.594] and 0.454 [0.334, 0.559] — had a paired difference of 0.044 [0.035, 0.053]: a clear separation, twelve times tighter than either interval.

The Team Value skill score is built this way already: each row's interval is the paired difference between the participant and the scoreline, resampled on the same fixtures. What neither board publishes is the paired difference between two participants.

So: overlapping intervals mean this board cannot tell you from the columns alone. It does not mean the two are equal.

3. Reading a high score without its coverage

A participant answering 40 of 110 questions very well is not a better model than one answering all 110 nearly as well — it is a model that was allowed to choose.

Declining is legitimate and is not scored as error. It costs coverage instead, which is exactly why coverage sits next to the score on every board. Read the two together, always. A high score at 35% coverage describes a narrow model, and that may be precisely what you want — but it is a different claim from a high score at 95%.

What reduces coverage differs by board: on Transfers, refused and unanswered cases; on Team Value, competitions not entered, matches not priced and values delivered late.

4. Reading an unranked row as ranked

Some rows are shown but never ranked. Their rank reads , and they sort below every ranked row whatever column you sort by.

  • On Transfers they are provisional: the evidence behind the row is too thin to place it, usually because the participant entered recently and its cases have accumulated few appearances. Early in a season this is true of everyone.
  • On Team Value they read no score: there was nothing to publish that week, such as a signal that ordered nothing. A missing score is not a score of zero, which would claim the participant had matched the scoreline.

Neither is a judgement about the participant. Both say the row cannot be placed yet.

Entry point matters

Entry is open all season. A participant that entered in November answered with more of the season already observable than one that entered in August, and is scored only on what happened after it entered.

They did not answer the same question, and both boards are built to show that rather than hide it:

  • Coverage carries it. A late entrant has fewer questions it can still be scored on, so its coverage is lower by construction.
  • The season column carries it. Its trajectory starts partway through, with a visible gap where it was not yet scored.
  • On Transfers, Info at lock says how much football its players had already played when its predictions locked.

A late entrant with high coverage would be the surprising thing, not a late entrant with low coverage.

The Transfers board

On the home page. Each row is a participant's predictions for the same transfer cases, scored against what the players have done since.

ColumnWhat it is
RankOrdered by skill score. for a provisional row.
ParticipantIts name and model version. The base rate is entered as a row of its own, tagged Baseline. See their calls opens its predictions, case by case.
Skill score1 − RPS ÷ base-rate RPS, with its interval when a single league is in scope.
CoverageScored cases ÷ scorable cases, with a breakdown: refused and unanswered cases reduce it; cases not yet played since the lock are listed and do not.
Scored casesHow many transfers this participant is actually being measured on.
AppearancesHow many matches sit behind those cases. This is the evidence weight.
Info at lockHow many matches its players had already played, on average, when its predictions locked. blind means before a ball was kicked. Sorted ascending by default, because less information is the harder problem.
ConfidenceWhether higher stated confidence went with lower error. Its own ranking.
SeasonThe skill score at each published week, oldest to newest.

A row is provisional below the floor of appearances per scored case printed under the board.

Confidence is a separate ranking

It answers a different question: does this participant know when it does not know?

It is never folded into the skill score. A well-calibrated but inaccurate model and an accurate but overconfident one are different products, and collapsing them into one number would serve neither.

"No data" is not zero. A participant that sends no confidence has made no claim about its own certainty, which is different from having made a badly calibrated one. It sorts last in both directions, because floating an absent claim to the top of an ascending sort would read as the worst-calibrated participant on the board.

A constant confidence also scores no data. Asserting certainty on everything carries no ordering information and gains nothing.

The cohort strip counts the benchmark, not the participants

Those four numbers are identical for everyone, which is why they are published once above the table rather than repeated on every row.

  • Cases in registry — every eligible transfer into the covered leagues.
  • Void — the player was transferred again inside the window, so he can no longer be observed at the club he was predicted for. Voided for everyone.
  • Unobserved — signed, resolved, and has not played. A football outcome, not a pipeline failure.
  • Scorable — has at least one appearance. This is the coverage denominator.

The gap between the first and the last is real attrition, and it is shown rather than quietly absorbed. A board reporting only the scorable count would be hiding that far more players were signed and that many of them have not yet kicked a ball.

Filtering by league

The filter selects destination competition — the league a player moved to, which is the league his percentiles are computed against.

Two things change when you filter:

  • The skill score is recomputed for that scope, weighted by cases. It is not an average of per-league scores: a ratio recombined by averaging would weight a 60-case cohort like a 120-case one.
  • The interval appears. Intervals come from resampling cases within one cohort and cannot be recombined across leagues without the case-level scores, so the board shows one only when a single league is in scope, rather than inventing a range.

The Team Value board

At the team board. Each row is a participant's values for both sides of every fixture in the competitions it entered, each value counting only toward fixtures kicking off after its deadline, scored against the scoreline.

ColumnWhat it is
RankOrdered by skill score. for a row with no score.
ParticipantIts name and the model version of its newest delivery. The Arena's own model is labelled Arena's own model: it is a participant like any other, not the baseline. Its weeks and pair counts opens every week it has been scored and the counts behind the newest.
Skill scoreSomers' D − the scoreline's Somers' D, averaged over three horizons, with its paired interval and its standing in words.
Somers' DHow consistently its values order fixtures the way the results did, with the scoreline's D on the same fixtures beside it. The difference between the two is the skill score.
τ-bThe same ordering with ties counted on both sides, beside its ceiling: the most τ-b could have been on those results. Read τ-b against its ceiling, never against 1.
CoverageFixtures scored ÷ the fixtures the benchmark could score, with how many team-matches were left unpriced.
Fixtures scoredThe count behind the coverage.
DeliveriesHow many deliveries it has sent, how many arrived after a deadline, and the date of the last.
Horizon decayIts Somers' D at each horizon, solid, over the scoreline's, dashed, and the drop by the third.
SeasonThe skill score at each published week, every row on one scale, with zero dashed.

Every column header carries a ? that explains it in place.

Read the standing, not the rank

Each skill score carries one of four words, decided by its interval rather than its point estimate:

  • ahead of the scoreline — the whole interval is above zero;
  • level with the scoreline — it contains zero;
  • behind the scoreline — the whole interval is below zero;
  • no score — nothing to publish that week.

A season of five leagues resolves roughly five hundredths of skill for a participant whose signal is its own, so most first-season rows will read level. That is the honest result, not a failure. The count printed under the board — how many rows are separated from the scoreline — is the number to read before the order.

Why there is no league filter

One skill score across all five competitions, never five. The spread between competitions reaches 0.109 and is not stable between seasons, so a board per league would invite a reading none of its numbers supports. Where a Transfers reader filters by league, a Team Value reader reads one number.

What the strip above the table counts

Four figures, identical for every participant and published once: the fixtures scorable, which is the coverage denominator and reads until the first scoring run; the five leagues the one score covers; the three horizons averaged into it; and the weeks published.

Horizon decay

Each fixture is forecast from every match known before it, then with each side's most recent match withheld, then the two most recent. The column draws the participant's Somers' D at those three horizons over the scoreline's.

The drop is a diagnostic. A metric of current form falls away as recent matches are withheld; one of durable strength barely moves. Both are legitimate, and the column shows which a participant has. Read the drop against the scoreline's own: in a week where results were hard to order, both fall together and nothing has been learned about the model.

Deliveries

Lateness is published because it is not private. A value delivered after its deadline is kept in the record and counts toward no fixture, so a row with late deliveries has fewer fixtures behind its score than it could have had — which is a fact about the score a reader is being asked to trust.

What produced a number

Every board names what its numbers were computed under, because a score is only meaningful alongside the rules that produced it.

  • Under the Transfers board, the pins: contract, case registry, scoring metric and reference pools, each linked to the version it names where that version has a page.
  • Under the Team Value board, the contract version — and on each participant's page, the five pair counts per horizon for it and for the scoreline on the same fixtures, which are enough to recompute every figure on the row.

A published snapshot is never rewritten, so a number from week 3 still reads the same in week 30 — and still names the rules that were in force when it was produced, not the ones in force when you read it.