Glossary

The terms a board uses, defined once. Several of them are ordinary words doing specific work, and reading them loosely is the fastest way to misread a score.

Shared terms mean the same thing on every board. Where benchmarks compute one differently, the definition says what it always means and then gives each benchmark's formula — so skill score is one idea with two formulas, not two ideas sharing a name. The terms after them belong to one benchmark.

Shared terms

Baseline

What a participant is measured against: the answer anyone could give without a model of their own. It is chosen to be reachable by anyone, so that zero is a fair floor rather than an insider's number.

  • Transfers — the league base rate for the player's role and metric. A band's width is its share of peer appearances, so predicting the widths of the grid you were sent is the baseline.
  • Game Value — the scoreline: the same forecast built from goal difference instead of values, over the same fixtures. A participant that copies the scoreline scores exactly zero.

Skill score

How much better a participant did than the baseline, on the same questions.

  • Zero — level with the baseline.
  • Above zero — the participant knew something the baseline did not. This is the whole claim.
  • Below zero — worse than the baseline. Published as negative, never clamped: a confidently wrong model belongs below a shrug.

Each benchmark computes it from its own statistic:

  • Transfers1 − participant RPS ÷ baseline RPS, over the participant's scored cases.
  • Game Valueparticipant Somers' D − baseline Somers' D, each averaged over the three horizons, over the participant's scored fixtures.

The two scales are not comparable. A score is read against its own board's zero, never against another benchmark's.

Interval

The 95% range around a published skill score, from resampling the questions it was scored on. Early in a season the interval is the story, and a board that showed only point estimates would invite a conclusion the data does not support.

The unit resampled is always the independent question, never its parts:

  • Transfers — cases, never appearances: appearances within a case share a player and are not independent evidence. Computed per league cohort, so a board shows one when a single league is in scope.
  • Game Value — fixtures, stratified within competition, never fixture-and-horizon pairs, which share a result. Paired: participant and baseline are scored on the same resampled fixtures, so their difference is resolved far better than either number alone.

Coverage

What a participant was scored on, over what it could have been scored on. The denominator is the same for every participant, so declining, gaps and arriving late are all visible and costed — without any of them being scored as a wrong answer. A participant answering a narrow slice very well is not a better model than one answering everything nearly as well.

  • Transfers — scored cases over scorable cases. A refusal and a case never answered both reduce it.
  • Game Value — fixtures scored over the fixtures the benchmark could score, across the whole season and all five competitions. A competition not entered, a match not priced and a value delivered late all reduce it.

Dry run

A submission put through every real check with nothing stored. It answers exactly as a real submission would, so passing one guarantees acceptance, and it counts toward no limit.

  • Transfers — optional: a verdict per case, run as often as you like before submitting for real.
  • Game Value — required once: certification is a dry run, and passing it is enrolment. See certification and enrolment.

Snapshot

One scoring run's published result, written once and never edited. Each weekly run is rebuilt from the whole season and appended as a new snapshot, so a value published in week 3 still reads the same in week 30, and a board can show a participant's trajectory rather than only its latest value.

Version pins

The versions a snapshot was computed under, recorded with it. They are what make a score citable weeks later: the rules that produced it are named and frozen, rather than being whatever the current documentation happens to say.

  • Transfers — contract, case registry, entity crosswalk, reference pool and scoring metric.
  • Game Value — contract and baseline version, with the week's bootstrap seed and replicate count, which reproduce its interval exactly.

Publication

A scored participant appearing on a public board. Scoring and publication are separate: a participant is scored whether or not it is published, and appears once a person has confirmed who it is. Publication is resolved when the board is read, so a late confirmation reveals the whole trajectory at once and costs the participant nothing.

A published participant's individual answers are published too. Transfers publishes them case by case; Game Value's match page, showing each vendor's values week by week, is planned and not yet built.

Transfers

Case

One transfer. A player, a move, and a destination — the unit a participant is asked about and the unit a score is built from.

A case is not a player: the same player transferred twice is two cases. It is not a match either; one case accumulates many appearances over a season, and all of them together produce the one realised distribution that case is scored against.

Cases are weighted equally in a participant's score, regardless of how many appearances each accumulated. An ever-present signing does not outweigh the rest of the cohort.

Appearance

One player in one match. Every appearance counts, with no minutes floor: playing time is already expressed in the percentile a short appearance lands on, and a substitute's twenty minutes produces a low percentile that is the correct observation rather than noise.

Values are raw per-match totals, never per-90 rates.

Cohort

Every case for one competition and season. A cohort's counts are properties of the benchmark, identical for every participant, and are published once rather than per row:

  • Candidate — cases in the registry.
  • Void — the player transferred again inside the observation window, so he can no longer be observed at the club he was predicted for. Voided for everyone.
  • Unobserved — signed, resolved, and has not played. A football outcome, not a pipeline failure.
  • Scorable — has recorded at least one appearance. This is the denominator for coverage.

Refusal

A participant declining a case, by marking it refused with a reason in its submission. All-or-nothing: every metric in the case is answered, or the case is refused.

A refusal is not scored as a maximally wrong prediction. A model that knows a case is outside its scope is behaving correctly, and scoring it as wrong would make good judgement indistinguishable from bad modelling. It costs coverage instead.

Realised distribution

What actually happened, as a distribution over the same bands the prediction used. Each appearance is banded, and the histogram of those bands across every appearance to date is what a prediction is scored against.

It is rebuilt from full history every week, not accumulated. So a score is always the whole season to date, and adding a week cannot corrupt what earlier weeks measured.

Ranked probability score

RPS — the distance between the predicted distribution and the realised one, comparing their cumulative distributions band by band. Bounded in [0, 1]; lower is better.

Two properties matter. It is order-aware, so being one band out costs less than being five bands out. And it is proper, meaning a participant minimises its expected score by reporting what it actually believes — there is no shading that improves the expected result.

Confidence discrimination

Whether a participant's stated confidence tracked its realised error — a rank correlation between confidence and -RPS across its scored cases.

Published as its own ranked column, never folded into the skill score. A participant that sends no confidence shows no data, never zero, and its accuracy ranking is untouched.

A constant confidence, high or low, carries no rank information and scores no data. Asserting certainty gains nothing.

Provisional

A row shown but not ranked, because it sits below the minimum number of observed appearances.

Provisional is not a judgement about the participant. It says the evidence is too thin to place the row, which early in a season is true of everyone.

Game Value — team

The team board's terms, in the order a participant meets them. The delivery contract and scoring have each one in full.

Fixture

One match, as the Arena publishes it: its own match_id (fam_2627_premierleague_0042), its round, kick-off, both teams, and the deadline its round is due at. A participant delivers one value for each side of a fixture, and both sides together.

Round

The matches of a competition's matchday, which share one deadline. A postponed match played more than five days from the rest of its round becomes a round of its own, with its own deadline and its own delivery budget.

Deadline

When a round's values are due: 23:59 Europe/Paris, a full day after the round's last match — Tuesday, for a weekend round. The day's delay exists because match data is not final at the whistle.

Freeze

What happens to a value at its deadline. Before it, a value can be replaced and the last one counts. After it, a first value still lands, marked late, and is never scored; a different value is refused as frozen.

A value counts only toward fixtures kicking off after its deadline. That is what stops the day's grace from being used to see the next result.

Certification and enrolment

A dry run of one completed round, both sides of every match, with nothing rejected — and the moment a participant is enrolled. Entries come back late on a completed round, which is expected.

The competitions certified are the ones scored, all or nothing within each. Matches that kicked off before certification are refused.

Receipt

The Arena's answer to a delivery, and the participant's proof of it: an outcome for every entry — accepted, replaced, unchanged, late, or rejected with a reason — the Arena's own receipt time, and a payload_sha256 of the body exactly as received.

Horizon

How much recent history is withheld when a fixture is forecast. Horizon 0 uses every match known before kick-off, horizon 1 withholds each side's most recent, horizon 2 the two most recent.

The skill score averages all three. The drop across them — horizon decay — separates a metric of current form, which falls away, from one of durable strength, which barely moves.

Burn-in

The history a fixture needs before it can be scored: at every horizon, both sides must have at least one match left after the withholding, and a rate built from matches the participant described. A fixture that falls short at any horizon is not scored at all, so all three horizons are measured on the same fixtures.

Somers' D

How consistently a signal orders fixtures the way their results did: (C − D) ÷ (n₀ − n_Y), from 1 for a perfect ordering to −1 for a perfectly backwards one. Pairs of fixtures the results themselves do not order — the same goal difference — are left out of the denominator, so they are held against nobody. It is the statistic the skill score is built from.

τ-b

The same concordance with ties kept on both sides: (C − D) ÷ √((n₀ − n_S) × (n₀ − n_Y)). It penalises a signal that hands many teams the same value, which Somers' D does not, and is published beside it so that being vague cannot look like being right.

Ceiling

The most τ-b could have been on a given set of results: √(1 − n_Y ÷ n₀). Pairs with the same goal difference stay in τ-b's denominator, so real football keeps it below 1, and the ceiling is what those results allowed. It belongs to τ-b, not to Somers' D, which can reach 1.

Pair counts

The five numbers every figure on the team board is computed from, per horizon: n₀ all pairs of scored fixtures, C pairs the signal and the result order the same way, D pairs they order oppositely, n_S pairs tied on the signal, and n_Y pairs tied on the result.

They are published for each participant and for the baseline on the same fixtures, so any score can be recomputed rather than argued about.

Standing

How a team board row reads against the scoreline — decided by its interval, not its point estimate: ahead of the scoreline when the whole interval is above zero, behind when it is below, level when it contains zero, and no score when there is nothing to publish, such as a signal that ordered nothing.

Level is the expected first-season result, not a failure: a season resolves roughly five hundredths of skill for a participant whose signal is its own.

Awarded

A result decided off the pitch, by ruling. An awarded fixture keeps its values in the record and is never scored: it is neither an outcome nor part of either team's rate, and does not count toward a team's known matches. Its counterpart ruling, score stands, confirms a flagged result and leaves it scored.