Method · Estimands
Methodology
Written for people who will try to break it — statisticians, coaches, analysts. Every public number on Scryglass is meant to survive that kind of reading.
Current public ladder
The published team snapshot uses a time-varying Bradley–Terry state-space filter over verified completed series. One immutable organization identity persists across events; event labels never create a second team. A forecast is emitted at the verified series start and the result enters the state only at verified completion.
Historical LTA source labels are retained for audit: LTA North maps to LCS, LTA South maps to CBLOL, and an unqualified LTA row is an Americas cross-region event. International events are classified from their source competition, so “LCK Road to MSI” remains LCK. Only series with verified Bo1, Bo3, or Bo5 format, contiguous maps, a compatible terminal score, and explicit completion provenance enter the rating. Ambiguous, tied, gapped, or incomplete groups remain quarantined.
When the artifact declares z = 1.64485, the adjusted value is the filtered mean minus that multiple of its diagonal Gaussian approximation: a normal-approximation 5th percentile used for conservative ordering. It is not presented as an empirically calibrated 95% coverage bound and is withheld when the row, metadata, and arithmetic do not agree. Uncertainty grows during inactivity.
Current affiliation comes only from the reviewed Riot tournament registry and matching row-level participation evidence. A last observed team or league cannot create current membership. The public current-membership view is Tier 1 only until an authoritative registry establishes lower-tier participation under the same contract.
Model changes require chronological holdouts, immutable prediction rows, log loss, Brier score, calibration, dependence-aware intervals, and rank stability. A challenger is not promoted merely because its point estimate is better; the current immutable validation artifact records the selection, untouched test, uncertainty interval, and promotion decision. The frozen 1,771-series test also reports calibration separately for Bo1, Bo3, and Bo5. Ratings are comparable only inside one connected historical comparison component. The design follows the dynamic Bradley–Terry and state-space literature: dynamic BT, state-space skill models, and whole-history rating.
On the immutable 1,630-series final test, the dynamic model scored 0.58491 log loss versus 0.61905 for the selected rolling-Elo benchmark. The paired difference was -0.03414, with a 95% circular moving-block bootstrap interval from -0.04701 to -0.02087. Its Brier score was 0.19948 and ten-bin ECE was 0.02920. The 221-series Bo5 slice had ECE 0.097, so this release is not described as universally calibrated or state of the art.
Sequential Dual Elo benchmark
The sequential Dual Elo track remains available as a time-safe pre-match feature benchmark. Each team carries a regional component and an international (meta) component. They sum to a total rating μ with a spread σ. Outcomes in the Oracle's Elixir pack update both. League chips on the Ratings page filter who appears while μ stays shared.
σ shrinks toward a floor as informative games arrive. Team floor is 25. When σ sits on that floor, the Evidence label reads Settled — the rating is as tight as this model allows. Headroom above the floor is what the adjusted rating penalizes:
benchmark adjusted rating = μ − max(0, σ − σ_min)
That soft penalty stops a thin regional spike from outranking a settled major org on the default ladder sort. Full LaTeX is in the formulas download.
Player Dual Elo
The legacy player track applies the same team outcome to the five observed players. It can travel across organizations and provide a shared lineup signal, but team wins generally do not identify individual teammate effects. Missing identifiability metadata is therefore “unverified,” never individual evidence. Exact shared-exposure cohorts remain tied, and public individual rank ordering is withheld even when a player has a unique exposure history.
A separately gated research model estimates role-relative 15-minute resource performance from gold, experience, and creep-score differentials without using map wins. It appears only when the current immutable pack contains a matching snapshot, metadata, and passing chronological validation artifact. Even then, it is not general skill, win contribution, or complete-game performance.
Draft Score
The published draft win probability is withheld by the current immutable validation artifact. On its untouched 2,491-map chronological test window, the best role/champion composition candidate scored 0.69427 log loss and 0.24997 Brier, compared with 0.69021 and 0.24853 for the overall blue-side base rate. Those scores belong to the pre-final fitted pipeline evaluated on that untouched window. The experimental Sandbox runtime was refit on the full population, including those labels, and is exposed only as a composition utility; it is not the exact coefficient artifact scored above. Calling either output a calibrated win chance would overstate the evidence.
For a complete, uniquely role-assigned draft, the frozen probability specification is logit(p_blue) = a + b × (β_side + e_composition). The composition edge contains role-aware direct terms, within-team pair terms, and all cross-team interactions. Its sign reverses when the two role-labelled compositions swap. The blue-side probability itself need not sum to one with that swapped counterfactual because the fitted side baseline remains attached to blue. The interval uses the model-intercept variance, a diagonal approximation for active composition terms, and the full intercept/slope covariance from the chronological calibration fit; omitted feature covariance remains a stated limitation.
The Draft Sandbox therefore uses a 0–100 experimental composition policy value at every draft state, including a complete board. It combines role-aware champion terms, observed ally pair terms, enemy interactions, and a bounded two-ply beam-minimax response search. It is useful for comparing branches inside this model, but it is not a win probability, an exhaustive best response, or evidence of optimal drafting. It also does not include private champion pools, scrim plans, or unannounced flex intent. Public patch labels are explicit contracts: public 25.xx maps to source key 15.xx and public 26.xx maps to source key 16.xx. Ambiguous one-digit minor strings are rejected.
Chronological kills benchmark
A match may be compared with the same-league mean from strictly earlier maps in the pack year. This is a simple chronological benchmark, not a fitted kills forecast. Missing history or map kills suppresses the comparison.
Evidence labels
Evidence semantics are model-specific. Dynamic team rows use the model-declared local spread and bound contract; Player Dual Elo uses its own sigma floor only after outcome identifiability is known. A shared cohort or missing metadata cannot be labelled as individual evidence merely because the numeric spread is small.
Blind / Counter (if shown)
Any Blind/Counter framing on Scryglass is an Oracle's Elixir matchup-shape proxy. It describes observed champion pairs; pick-order seats are outside the estimand.
Pack years
Default pack years are 2025–2026. Column-trimmed OE parquet, rating snapshots, and pinned calibrations live under /packs/. Cite the pack id when matching a published finding. Rebuild notes and essentials are on Reproduce.
Freshness and sources
Scryglass publishes versioned packs, with each rating file preserved as a dated snapshot. OE is the canonical inclusion source when the same game appears in both feeds. Verified GRID results may bridge a completed-game gap while the next OE export is pending; GRID may also provide event detail for a map whose canonical result is already OE-backed. Those roles are published separately as canonical_map_source and map_detail_source. Scheduled or scrim-like series stay outside the result set.
A new pack is built by the refresh workflow, then the public pointer is updated. The pack date is publication time; source metadata identifies the newest match. Use the details in Reproduce when the distinction matters.
Void grubs
The void-grubs article leads with a Patch 26.11+ opportunity-cost sensitivity. At even gold, leaving for the two-wave farm reference is worth more than contesting until estimated fight-win reaches its 58.24% contest bar. Gold@10 → map-win is an associational logit conversion; the result is not an identified action policy.
FAQ
Why can two teammates have the same score? When they have the same signed map exposure, team outcomes contain no information that separates them. The interface keeps that tie and states the identifiability limit.
Do league chips change Elo? They filter the ladder roster while the shared Elo stays fixed.
Where is probability quality? Favorite hit rate is only a threshold diagnostic. Release evidence uses chronological log loss, Brier score, calibration, and dependence-aware comparisons tied to an immutable model and pack.
Changelog
# Scryglass changelog ## 2026-07-25 - Public Dual Elo ladders with Trust (Settled / Thin), league chips, player profiles. - Matches gallery groups Bo-series; Dual Elo year accuracy on the Matches page. - H2H series boards with shareable URLs; model checklist stays on match pages. - Reproduce lists curated essentials only. - Draft WR API runs in TypeScript (no Python on Vercel).