Evaluation protocol

The rules that hold across every benchmark. A task decides what is measured; this page decides how any measurement here is conducted, published, and allowed to change.

Each rule is stated once, for every live benchmark. Where the benchmarks implement one differently, the table at the end says how, and each benchmark's own contract has the detail.

Capabilities, not companies

There is no universal benchmark, because there is no universal task. Valuing what a team did in a match is not projecting a transfer. Projecting a transfer is not player similarity. Player similarity is not opponent analysis. Each deserves its own methodology, and a single vendor score would average away the only thing a club actually wants to know.

So a participant does not receive a score. It receives one per benchmark, and a provider strong at projecting transfers and weak at valuing team performances reads as exactly that.

Six principles

Task-first. Every benchmark is designed around one question with a checkable answer, rather than a shared scoring template stretched to fit.

Use the strongest available signal. Where ground truth exists — what a player went on to do, what a match actually produced — it is used. Where it does not, the benchmark says so rather than substituting a proxy and calling it truth.

Neutrality. Every participant is evaluated under exactly the same protocol, against the same contract, on the same cases or fixtures. The Arena's own models take part on those terms and have no privileged access to reference data.

Transparency. Every benchmark states what is measured, why, how the score is computed, and where the data comes from. Every published result is reproducible from stored artifacts.

Versioning. Protocols, datasets and metrics carry versions. Every published score names the versions it was computed under, so historical results stay readable while the methodology improves.

Continuous evolution. Football changes and so does analytics. A frozen protocol would be accurate for one season and misleading afterwards.

Participation rules

Entering

Every benchmark is entered the same way.

  1. Register an identity. One registration and one API key for the whole Arena, whichever benchmarks you enter. The key is shown once.
  2. Enter a benchmark. Registering says who you are; entering says what you are competing in.
  3. Check the integration with a dry run. A dry run puts a submission through every real check and stores nothing, so passing it guarantees the real one is accepted. It can be run as often as needed. Where a benchmark requires one before anything counts, the table below says so.

Participation is by submission: a participant sends its answers to the Arena, and the Arena calls nothing. There is no endpoint to expose and no service to keep running, so a participant's infrastructure can never become part of its score.

Entry is open at any point in a season. See rolling entry.

What is pinned

Everything a score depends on is pinned in a contract version before it counts: the payload, the scoring rules, and whatever the benchmark scores against. A change to any of them is a new contract version, and a submission built against a different version is refused whole, with nothing stored, rather than scored under rules it was not written for.

What a benchmark scores against is fixed too, so a published score cannot be rewritten underneath it. What is deliberately not fixed is anything the football itself changes: a fixture list moves when matches are postponed, and says so as it happens.

When an answer is fixed

The rule every benchmark enforces: nothing counts toward an outcome that could have been known when the answer was fixed. A benchmark that let a participant answer after watching the result would be measuring memory, not models.

How an answer is fixed depends on what is being forecast, and the table below gives each benchmark's mechanism. Whatever the mechanism:

  • model_version is recorded with every answer and published with the score, so a board row is always traceable to the model that produced it and a later model cannot claim an earlier result.
  • Answers are keyed on the Arena's own identifiers, which carry nothing about where the underlying data came from, so an answer cannot be keyed against a scraped copy of the result.

Rolling entry

A benchmark that only accepted entrants in August would be a benchmark almost nobody could enter.

A participant entering mid-season answers from that moment, using everything observable by then, and is scored only on what happens afterwards. It is never scored on football it had already watched.

What arriving late costs is coverage and evidence, not accuracy: fewer questions answered, fewer outcomes behind each score, and a wider interval. Two participants that entered at different points did not answer the same question, and the board is built to show that rather than hide it.

Declining

A participant may leave questions unanswered, and declining is never scored as a wrong answer. A model that knows a question is outside its scope is behaving correctly, and scoring it as wrong would make good judgement indistinguishable from bad modelling.

What declining costs is coverage, published beside the score. Each benchmark sets the grain at which a participant may decline — a single case, or a whole competition — and states it in its contract.

Invalid submissions

A submission that breaks the contract is not a wrong answer either. Entries that fail validation are rejected with a reason and not stored: they are neither scored nor counted as declined, and they can be sent again once fixed. One bad entry never costs the rest of a submission.

Publication policy

Scores are appended, never rewritten. Each weekly run is computed afresh from the whole season, and the result is a new immutable snapshot. A value published in week 3 still reads the same in week 30.

That is what allows a board to show a participant's trajectory across a season rather than only its current value, and what makes a score citable weeks after it was produced.

Every snapshot names its versions — the contract it was scored under and every versioned input it used — so a score leads to the rules that produced it: not the current documentation, those rules, as of that week.

Uncertainty is published with the number. Every score is computed with an interval, and the board shows it wherever it can be computed honestly. Early in a season the interval is the story, and a point estimate alone would invite a conclusion the data does not support.

No privileged view of results. There is no vendor-only board. Every published score, and the evidence behind it, is on the same public pages for everyone, the participant included. What is not public is narrow and named: a participant's contact address; any call that answers only for the key that makes it, such as what a participant still owes; and a scored participant until a person has confirmed who it is. Confirmation reveals the participant's whole trajectory at once, so confirming late costs it nothing.

Individual answers are published, not only the aggregate. A participant's answer to a single question — one case, one match — is published under its name, beside every other participant's answer to the same question.

This follows from what a benchmark is. A score nobody can inspect is a claim, not a measurement: a reader who cannot see the individual answers has to take the aggregate on trust, which is the position the Arena exists to make unnecessary. It is also the only way a participant can be defended against a bad-looking number, because the question, the evidence and the baseline are all visible next to it.

An individual answer is published with what it takes to read it fairly: how much evidence it has been scored on, because no single answer is evidence about a model, and how much was already known when it was fixed, because a participant that answered in November answered an easier question than one that answered blind in August.

A result below a publication floor is marked, not hidden. Where a benchmark sets a floor — the evidence a cohort needs before its scores mean anything — a result below it is published and says so.

What a benchmark must document

Before a benchmark accepts entries it publishes: the question, the submission contract, the scoring formula with worked examples, the data it is built on with its gaps stated and dated, and the versions in force. Where taking part asks more than a single submission, it also publishes the path through it, step by step.

If something is not decided, the documentation says it is open rather than implying it is settled.

How each live benchmark applies them

RuleTransfersGame Value — team
What is answeredA case: one transfer, a distribution per metricA match: one value per team, both sides together
Dry runA check, run as often as neededRequired: certifying on one completed round per competition is the moment of enrolment
When an answer is fixedLocked when stored, before the case's first observation; one answer per case, never overwrittenReplaceable until the round's deadline and frozen at it; counts only toward fixtures kicking off after that deadline
Rolling entryOffered the cases still open; scored only on appearances after the lockScored from certification; matches that kicked off before it are refused
DecliningPer case, marked refused with a reasonPer competition: within one entered, every match is delivered
Invalid entriesRejected with a reason; the case stays openRejected with a reason; three stored deliveries per round
What is fixed before the seasonReference pools, never recomputed; the case registry is versionedThe deadline rule, burn-in, horizons and statistic, all part of the contract
What each snapshot recordsContract, case registry, entity crosswalk, reference pool, scoring metricContract, baseline version, bootstrap seed and replicates
IntervalPer league cohortPaired, across all five competitions
Publication floorOne scorable case per club, on average, per competition-seasonNone beyond the burn-in
Individual answersPublished, case by caseRecorded; a match page publishing each vendor's values week by week is planned and not yet built
The detailSubmission contractOnboarding and delivery contract