AInsights

AI Insights Analysis

Full Ranking Sources

Methodology · Mixed Core 07

AInsights Index Methodology

AIndex Mixed Core 07 uses three boards with two Core items each and two boards with a single Core item. A dual-Core item supplies one half of its board base; a single Core supplies the full base. Reviewed Extra Tests can add capped positive-residual evidence. The final score uses fixed board weights of 24%, 24%, 27%, 16%, and 9%.

Default Ranking at a Glance

  1. Canonicalize source-backed, direct 0–100 results and select the default deduplicated model population before fitting.
  2. Exclude a model when any board has no observed Core, or at least four of the eight configured Core items are missing. Sum the fixed Core shares without renormalizing missing items.
  3. On each board, use only models with all its configured Core observed to fit anonymous extension OLS trends and award extension bonuses.
  4. Keep only positive residuals, combine them with a zero-neutral monotone log-sum-exp function, and cap the increment using the pooled positive-residual mean + sqrt(2) × SD.
  5. Calculate board_score = min(100, core_score + capped_bonus), then combine the five board scores with their fixed weights and sort the unrounded final score descending.

Capability Boards

The published weights remain fixed between data refreshes. All eight Core tests belong to distinct benchmark families.

Board Share Configured Core Published score
Coding 24% Terminal-Bench v4.0 + SciCode coding_score
Agentic/tool work 24% AutomationBench-AA + τ³-Banking agentic-tool-work_score
Hard reasoning 27% CritPt (single Core) hard-reasoning_score
Knowledge/science 16% AA-Omniscience Accuracy + GDP.pdf knowledge-science_score
Instruction/context 9% AA-LCR v1.1 (single Core) instruction-context_score

The Core slots cover five capabilities with distinct task families and direct bounded scores. Each canonical family occupies one Core slot across the scheme and cannot also supply Extra Test evidence. The Terminal-Bench Core slot uses the v4.0 release and its matching score protocol.

Core Score

Each configured Core item retains a fixed share of its board base:

single_core_score = adjusted_score; dual_core_score = adjusted_1 / 2 + adjusted_2 / 2

Core uses direct bounded 0–100 values. A missing Core share contributes zero points but remains missing evidence; an observed numerical zero is valid evidence. The available half of a dual-Core board is never increased to compensate for a missing half. Unconfigured second slots do not count as missing.

A model is not scored when any board has no observed Core, or when at least four configured Core items are missing. A single-Core board therefore requires its one observation; a dual-Core board requires at least one. Core data is never borrowed from a sibling model or another effort configuration.

Selecting Extra Tests

Extra Tests target demanding tasks that still distinguish frontier models. We review the score distribution for ceiling concentration, the challenge of the tasks, and whether the metric measures a capability that the board is intended to represent.

The benchmark controller must be independent of ranked model vendors. A public owner or evaluator source must document the task set, scoring rule, and result protocol. We pin the benchmark release, metric, tool setting, and agent harness where applicable, and disclose material protocol limits with the result. Different versions or scoring protocols receive separate identities. Equivalent representations of one canonical family cannot create duplicate bonus opportunities inside a board.

A test needs at least three exact-configuration default representatives with complete Core evidence on its board to fit a trend. Smaller cohorts can qualify under this rule; their observation counts and fitted parameters remain visible so readers can judge how sensitive the estimate is to new results.

Extension evidence is only-add. Missing extension results remain absent, contribute no residual, receive no imputed zero or neutral score, and never reduce Core. A board receives no extension bonus while any of its configured Core items are missing; broader extensions cannot repair a missing Core share.

Which Extra Tests Can Add Points?

The registry lists each eligible Extra Test with its board, release, score metric, and result protocol. Open a test to inspect its scores and source details.

BoardExtension testResult source and scope
CodingFrontierCode v1.1 MainCognition official Main leaderboard; model-specific agent harnesses are disclosed. Cognition also builds agents.
CodingDeepSWE v1.1 owner runDataCurve 113-task set, mini-swe-agent, four runs per task; scored-attempt pass@1.
CodingSWE-Marathon v1.1 owner runOwner-run set of 20 long software tasks with eight binary trials each; agent harnesses vary by model.
Agentic/tool workAPEX-Agents-AAAA runs 452 Mercor tasks with one Stirrup agent, Archipelago MCP environment, three repeats and strict pass@1. Task and agent commits are not publicly pinned.
Agentic/tool workAA-AnalystAgent pass^5AA private 80-task analyst evaluation; all five independent attempts must pass.
Agentic/tool workTerminal-Bench-Science v0.1.0AA mini-swe-agent, mean pass@1 over three runs on 70 scientific tasks.
Agentic/tool workAA-Briefcase v1.1 Rubric Pass RateAA private 91-task document and artifact evaluation; native rubric pass fraction.
Agentic/tool workHarvey LAB-AA All-pass RateAA private 120-task legal evaluation; all rubric criteria must pass. A single LLM judge grades the rubric.
Agentic/tool workToolathlon-Verified official Pass@1HKUST owner leaderboard with the Default agent and Pass@1 scoring.
Agentic/tool workOSWorld 2.0 v2026.06.24Official full 108-task release, standard tools, 500 steps, binary completion.
Agentic/tool workARC-AGI-3 StandardARC Prize evaluation with the Standard harness.
Agentic/tool workAgents' Last Exam V1 Overall pass rateBerkeley RDI owner-run pass rate; model-specific agent harnesses are disclosed.
Hard reasoningFrontierMath Tier 4 v2Epoch-run official Tier 4 v2 export. OpenAI commissioned the benchmark and had problem access.
Hard reasoningChess Puzzles v1.1.6Epoch runs on 100 unpublished generated positions, given as FEN text with one best next move.
Hard reasoningMystery Game Puzzles v1.0.4Epoch runs on 100 hidden-identity games, text state and minimal agent scaffold; one best next move.
Hard reasoningEBR-bench v4 Card-banEpoch single-agent runs with the v4 Card-ban scoring rule.
Hard reasoningHumanity's Last ExamAA common-protocol evaluation of the no-tools test.
Knowledge/scienceMLCR-AA OverallAA private medical long-context tasks; Overall score combines accuracy, completeness, and conciseness.
Instruction/contextIFBench (AA)AA evaluates 294 questions over five repeats with the official loose mode; its Score is prompt-level accuracy.

Core task families cannot appear in the extension pool. Terminal-Bench-Science uses a distinct scientific task set, although its shared ecosystem and agent may correlate with Terminal-Bench v4 Core. New benchmark releases and scoring protocols undergo review before they can enter the registry. An Extra Test adds points only when an exact-configuration result has complete Core on its board and a positive residual above the fitted expectation.

Matching a Result to a Model

Each result is matched to its reported model checkpoint and effort configuration, as well as the test release, task set, scoring metric, tool setting, and agent harness. A result reported for one effort tier supplies evidence for that tier only. We retain the original source and protocol details so readers can trace each match.

After matching, the model must have every configured Core item on the Extra Test's board. The fitted trend gives the expected test score at that Core level. Only the amount by which the observed result exceeds that expectation enters the board bonus.

Anonymous Positive Residuals

For each extension item j, the current eligible deduplicated cohort supplies observed pairs from models with every Core item complete on the board: board Core score c_i and extension result y_ij. Mixed Core 07 fits:

predicted_ij = clip(alpha_j + max(0, beta_j) × c_i, 0, 100)

positive_residual_ij = max(y_ij - predicted_ij, 0)

The OLS intercept and non-negative slope are learned anonymously from the observed cohort with complete Core on the relevant board. A trend requires at least three observed representatives. Small cohorts make the fitted trend more sensitive to new results; counts and parameters are published for every test. They are benchmark-level calibration parameters, not corrections for a named model. An extension result at or below its cohort expectation contributes zero rather than a penalty.

Monotone Log-Sum-Exp Bonus and Dynamic Cap

For a model's positive observed residuals in one board, the raw bonus is:

raw_bonus = log(1 + sum(expm1(positive_residual_j)))

At fixed calibration this form is zero-neutral and monotone: adding a zero residual changes nothing, while adding any positive residual cannot lower the bonus. It allows multiple independent frontier signals to corroborate one another.

One symmetric cap is recalculated from all pooled positive extension residuals in the current default cohort:

cap = mean(pooled_positive_residuals) + sqrt(2) × population_SD(pooled_positive_residuals)

bonus = min(cap, raw_bonus)

The cap is calculated from the observed cohort and can change after a data refresh. Capping limits repeated corroboration, but models with broader observed extension coverage still have more opportunities to produce a positive residual; extension counts are therefore published alongside scores.

Board and Final Score

board_score_b = min(100, core_score_b + bonus_b)

board_points_b = board_score_b × board_weight_b / 100

AIndex = 0.24 × coding + 0.24 × agentic + 0.27 × reasoning + 0.16 × knowledge + 0.09 × context

The displayed score, bar length, and rank all use this same quantity. Ranking follows the unrounded AIndex value; stable model ID resolves an exact numerical tie.

Deduplication and Exact Configurations

The default leaderboard first selects one representative exact configuration for each variantGroup by descending variantPriority and ascending slug, before checking eligibility. It never switches to a lower configuration to repair missing Core. Anonymous extension trends and the dynamic cap are then fitted on the eligible deduplicated cohort. This is also the calibration population used by the published ranking.

When Full Ranking shows every eligible tier, exact configurations use their own Core results and reuse the same serialized deduplicated trends and cap. Only external evidence explicitly scoped to that exact variant with variantScoped may contribute to its extension bonus; family-level evidence is not concentrated onto one effort tier. The production export records the ranking grain and selected exact slug so the join can be audited.

Radar and Evidence Coverage

The first five radar axes read the five Mixed Core 07 board scores directly. The sixth shows the selected visual benchmark, with AA MMMU-Pro as the default. Other verified visual tests remain separate by version and protocol; missing results are not plotted as zero. Changing this visual selection does not change AIndex.

Core and extension coverage are disclosed alongside the board scores. Missing Core shares contribute no points and can prevent ranking or board extensions. Missing extension cells remain absent.

Sensitivity Views and Custom Tools

Equal-board 2PL, Core Rasch, Sparse Rasch, and Dense Rasch remain available as sensitivity views. They do not contribute a fixed percentage to Mixed Core 07 and cannot override its score order. Full Ranking may display standalone 2PL and Dense Rasch ranks for comparison only.

Custom tools can separately explore method ranks, capability-board weights, and per-benchmark calculations. These user-defined views do not reproduce or modify the default AIndex unless they implement the complete Mixed Core 07 Core, residual, cap, and five-board pipeline.

AIndex Calculation

  1. Load source-backed direct scores and canonical benchmark policy metadata.
  2. Fix the deduplicated representatives before checking for any empty Core board or at least four missing configured items.
  3. Sum the fixed Core shares: one full share for a single-Core board, two half shares for a dual-Core board.
  4. On complete-Core boards only, fit non-negative-slope OLS trends and retain observed positive extension residuals.
  5. Pool positive residuals to derive the dynamic mean + sqrt(2) × SD cap.
  6. Apply the monotone log-sum-exp bonus, cap it, and add it to the matching Core without imputing missing extension results.
  7. Apply fixed board weights 24/24/27/16/9, sort the unrounded AIndex score descending, and publish Core, bonus, coverage, and sensitivity diagnostics.