Default Ranking at a Glance
- Canonicalize source-backed, direct 0–100 results and select the default deduplicated model population before fitting.
- Exclude a model when any board has no observed Core, or at least four of the eight configured Core items are missing. Sum the fixed Core shares without renormalizing missing items.
- On each board, use only models with all its configured Core observed to fit anonymous extension OLS trends and award extension bonuses.
- Keep only positive residuals, combine them with a zero-neutral monotone log-sum-exp function, and cap the increment using the pooled positive-residual
mean + sqrt(2) × SD.
- Calculate
board_score = min(100, core_score + capped_bonus), then combine the five board scores with their fixed weights and sort the unrounded final score descending.
Capability Boards
The published weights remain fixed between data refreshes. All eight Core tests belong to distinct benchmark families.
| Board |
Share |
Configured Core |
Published score |
| Coding |
24% |
Terminal-Bench v4.0 + SciCode |
coding_score |
| Agentic/tool work |
24% |
AutomationBench-AA + τ³-Banking |
agentic-tool-work_score |
| Hard reasoning |
27% |
CritPt (single Core) |
hard-reasoning_score |
| Knowledge/science |
16% |
AA-Omniscience Accuracy + GDP.pdf |
knowledge-science_score |
| Instruction/context |
9% |
AA-LCR v1.1 (single Core) |
instruction-context_score |
The Core slots cover five capabilities with distinct task families and direct bounded scores. Each canonical family occupies one Core slot across the scheme and cannot also supply Extra Test evidence. The Terminal-Bench Core slot uses the v4.0 release and its matching score protocol.
Core Score
Each configured Core item retains a fixed share of its board base:
single_core_score = adjusted_score; dual_core_score = adjusted_1 / 2 + adjusted_2 / 2
Core uses direct bounded 0–100 values. A missing Core share contributes zero points but remains missing evidence; an observed numerical zero is valid evidence. The available half of a dual-Core board is never increased to compensate for a missing half. Unconfigured second slots do not count as missing.
A model is not scored when any board has no observed Core, or when at least four configured Core items are missing. A single-Core board therefore requires its one observation; a dual-Core board requires at least one. Core data is never borrowed from a sibling model or another effort configuration.
Selecting Extra Tests
Extra Tests target demanding tasks that still distinguish frontier models. We review the score distribution for ceiling concentration, the challenge of the tasks, and whether the metric measures a capability that the board is intended to represent.
The benchmark controller must be independent of ranked model vendors. A public owner or evaluator source must document the task set, scoring rule, and result protocol. We pin the benchmark release, metric, tool setting, and agent harness where applicable, and disclose material protocol limits with the result. Different versions or scoring protocols receive separate identities. Equivalent representations of one canonical family cannot create duplicate bonus opportunities inside a board.
A test needs at least three exact-configuration default representatives with complete Core evidence on its board to fit a trend. Smaller cohorts can qualify under this rule; their observation counts and fitted parameters remain visible so readers can judge how sensitive the estimate is to new results.
Extension evidence is only-add. Missing extension results remain absent, contribute no residual, receive no imputed zero or neutral score, and never reduce Core. A board receives no extension bonus while any of its configured Core items are missing; broader extensions cannot repair a missing Core share.
Which Extra Tests Can Add Points?
The registry lists each eligible Extra Test with its board, release, score metric, and result protocol. Open a test to inspect its scores and source details.
| Board | Extension test | Result source and scope |
| Coding | FrontierCode v1.1 Main | Cognition official Main leaderboard; model-specific agent harnesses are disclosed. Cognition also builds agents. |
| Coding | DeepSWE v1.1 owner run | DataCurve 113-task set, mini-swe-agent, four runs per task; scored-attempt pass@1. |
| Coding | SWE-Marathon v1.1 owner run | Owner-run set of 20 long software tasks with eight binary trials each; agent harnesses vary by model. |
| Agentic/tool work | APEX-Agents-AA | AA runs 452 Mercor tasks with one Stirrup agent, Archipelago MCP environment, three repeats and strict pass@1. Task and agent commits are not publicly pinned. |
| Agentic/tool work | AA-AnalystAgent pass^5 | AA private 80-task analyst evaluation; all five independent attempts must pass. |
| Agentic/tool work | Terminal-Bench-Science v0.1.0 | AA mini-swe-agent, mean pass@1 over three runs on 70 scientific tasks. |
| Agentic/tool work | AA-Briefcase v1.1 Rubric Pass Rate | AA private 91-task document and artifact evaluation; native rubric pass fraction. |
| Agentic/tool work | Harvey LAB-AA All-pass Rate | AA private 120-task legal evaluation; all rubric criteria must pass. A single LLM judge grades the rubric. |
| Agentic/tool work | Toolathlon-Verified official Pass@1 | HKUST owner leaderboard with the Default agent and Pass@1 scoring. |
| Agentic/tool work | OSWorld 2.0 v2026.06.24 | Official full 108-task release, standard tools, 500 steps, binary completion. |
| Agentic/tool work | ARC-AGI-3 Standard | ARC Prize evaluation with the Standard harness. |
| Agentic/tool work | Agents' Last Exam V1 Overall pass rate | Berkeley RDI owner-run pass rate; model-specific agent harnesses are disclosed. |
| Hard reasoning | FrontierMath Tier 4 v2 | Epoch-run official Tier 4 v2 export. OpenAI commissioned the benchmark and had problem access. |
| Hard reasoning | Chess Puzzles v1.1.6 | Epoch runs on 100 unpublished generated positions, given as FEN text with one best next move. |
| Hard reasoning | Mystery Game Puzzles v1.0.4 | Epoch runs on 100 hidden-identity games, text state and minimal agent scaffold; one best next move. |
| Hard reasoning | EBR-bench v4 Card-ban | Epoch single-agent runs with the v4 Card-ban scoring rule. |
| Hard reasoning | Humanity's Last Exam | AA common-protocol evaluation of the no-tools test. |
| Knowledge/science | MLCR-AA Overall | AA private medical long-context tasks; Overall score combines accuracy, completeness, and conciseness. |
| Instruction/context | IFBench (AA) | AA evaluates 294 questions over five repeats with the official loose mode; its Score is prompt-level accuracy. |
Core task families cannot appear in the extension pool. Terminal-Bench-Science uses a distinct scientific task set, although its shared ecosystem and agent may correlate with Terminal-Bench v4 Core. New benchmark releases and scoring protocols undergo review before they can enter the registry. An Extra Test adds points only when an exact-configuration result has complete Core on its board and a positive residual above the fitted expectation.
Matching a Result to a Model
Each result is matched to its reported model checkpoint and effort configuration, as well as the test release, task set, scoring metric, tool setting, and agent harness. A result reported for one effort tier supplies evidence for that tier only. We retain the original source and protocol details so readers can trace each match.
After matching, the model must have every configured Core item on the Extra Test's board. The fitted trend gives the expected test score at that Core level. Only the amount by which the observed result exceeds that expectation enters the board bonus.
Anonymous Positive Residuals
For each extension item j, the current eligible deduplicated cohort supplies observed pairs from models with every Core item complete on the board: board Core score c_i and extension result y_ij. Mixed Core 07 fits:
predicted_ij = clip(alpha_j + max(0, beta_j) × c_i, 0, 100)
positive_residual_ij = max(y_ij - predicted_ij, 0)
The OLS intercept and non-negative slope are learned anonymously from the observed cohort with complete Core on the relevant board. A trend requires at least three observed representatives. Small cohorts make the fitted trend more sensitive to new results; counts and parameters are published for every test. They are benchmark-level calibration parameters, not corrections for a named model. An extension result at or below its cohort expectation contributes zero rather than a penalty.
Monotone Log-Sum-Exp Bonus and Dynamic Cap
For a model's positive observed residuals in one board, the raw bonus is:
raw_bonus = log(1 + sum(expm1(positive_residual_j)))
At fixed calibration this form is zero-neutral and monotone: adding a zero residual changes nothing, while adding any positive residual cannot lower the bonus. It allows multiple independent frontier signals to corroborate one another.
One symmetric cap is recalculated from all pooled positive extension residuals in the current default cohort:
cap = mean(pooled_positive_residuals) + sqrt(2) × population_SD(pooled_positive_residuals)
bonus = min(cap, raw_bonus)
The cap is calculated from the observed cohort and can change after a data refresh. Capping limits repeated corroboration, but models with broader observed extension coverage still have more opportunities to produce a positive residual; extension counts are therefore published alongside scores.
Board and Final Score
board_score_b = min(100, core_score_b + bonus_b)
board_points_b = board_score_b × board_weight_b / 100
AIndex = 0.24 × coding + 0.24 × agentic + 0.27 × reasoning + 0.16 × knowledge + 0.09 × context
The displayed score, bar length, and rank all use this same quantity. Ranking follows the unrounded AIndex value; stable model ID resolves an exact numerical tie.
Deduplication and Exact Configurations
The default leaderboard first selects one representative exact configuration for each variantGroup by descending variantPriority and ascending slug, before checking eligibility. It never switches to a lower configuration to repair missing Core. Anonymous extension trends and the dynamic cap are then fitted on the eligible deduplicated cohort. This is also the calibration population used by the published ranking.
When Full Ranking shows every eligible tier, exact configurations use their own Core results and reuse the same serialized deduplicated trends and cap. Only external evidence explicitly scoped to that exact variant with variantScoped may contribute to its extension bonus; family-level evidence is not concentrated onto one effort tier. The production export records the ranking grain and selected exact slug so the join can be audited.
Radar and Evidence Coverage
The first five radar axes read the five Mixed Core 07 board scores directly. The sixth shows the selected visual benchmark, with AA MMMU-Pro as the default. Other verified visual tests remain separate by version and protocol; missing results are not plotted as zero. Changing this visual selection does not change AIndex.
Core and extension coverage are disclosed alongside the board scores. Missing Core shares contribute no points and can prevent ranking or board extensions. Missing extension cells remain absent.
Sensitivity Views and Custom Tools
Equal-board 2PL, Core Rasch, Sparse Rasch, and Dense Rasch remain available as sensitivity views. They do not contribute a fixed percentage to Mixed Core 07 and cannot override its score order. Full Ranking may display standalone 2PL and Dense Rasch ranks for comparison only.
Custom tools can separately explore method ranks, capability-board weights, and per-benchmark calculations. These user-defined views do not reproduce or modify the default AIndex unless they implement the complete Mixed Core 07 Core, residual, cap, and five-board pipeline.
AIndex Calculation
- Load source-backed direct scores and canonical benchmark policy metadata.
- Fix the deduplicated representatives before checking for any empty Core board or at least four missing configured items.
- Sum the fixed Core shares: one full share for a single-Core board, two half shares for a dual-Core board.
- On complete-Core boards only, fit non-negative-slope OLS trends and retain observed positive extension residuals.
- Pool positive residuals to derive the dynamic
mean + sqrt(2) × SD cap.
- Apply the monotone log-sum-exp bonus, cap it, and add it to the matching Core without imputing missing extension results.
- Apply fixed board weights 24/24/27/16/9, sort the unrounded AIndex score descending, and publish Core, bonus, coverage, and sensitivity diagnostics.