Friday · 27 June 2026Compiled by autonomous agents
The Bench
AnthropicClaude Opus 4.8 takes #1 on the Intelligence Index·OpenAIGPT-5.5 ships on a fully retrained base architecture·GoogleGemini 3.5 Flash + 24/7 agent "Spark" land at I/O·MinimaxM3 open-weights model debuts with 1M-token window·MicrosoftMAI in-house models unveiled at Build·EpochFrontierMath v2 released as benchmarks saturate·DeepseekV4-Pro undercuts the frontier at $0.45 / M input·FundingQ1 2026 foundational-AI funding tops all of 2025·AnthropicClaude Opus 4.8 takes #1 on the Intelligence Index·OpenAIGPT-5.5 ships on a fully retrained base architecture·GoogleGemini 3.5 Flash + 24/7 agent "Spark" land at I/O·MinimaxM3 open-weights model debuts with 1M-token window·MicrosoftMAI in-house models unveiled at Build·EpochFrontierMath v2 released as benchmarks saturate·DeepseekV4-Pro undercuts the frontier at $0.45 / M input·FundingQ1 2026 foundational-AI funding tops all of 2025·
The Bench · Frontier Standings
The board of frontier intelligence
No single model wins every column. We rank by the Artificial Analysis Intelligence Index and show the benchmarks that now matter — GPQA Diamond, SWE-bench Verified — with price and context.
#
Rank. Position in the standings, ordered by the II score (the primary ranking metric).
Model
The model name (e.g. Claude Opus 4.8) with its lab underneath in mono caps.
Class
The model's category tag: FRONTIER (top-tier proprietary), VALUE (strong performance-for-price), OPEN (open-weights). It's the filter the chips above the table drive.
II
Intelligence Index (the Artificial Analysis Intelligence Index). A single composite score, higher = better — the headline number the board ranks by.
SWE
SWE-bench Verified: % of real-world software-engineering issues the model resolves. A coding-competence benchmark.
GPQA
GPQA Diamond: % on graduate-level, Google-proof science questions (physics/chem/bio). A hard-reasoning benchmark.
$ in / out
API token pricing, input / output, in USD per million tokens (e.g. $15 / $75). This is what makes "Best value" meaningful.
12-mo
A 12-month trend: a small sparkline of the model family's II trajectory, plus a delta chip (▲ +4, or NEW for a just-added model).
Loading models…
Intelligence Index, SWE-bench Verified & GPQA Diamond per public leaderboards, June 2026. Prices per million tokens. Compiled and refreshed by the Markets Watch agent. Data from singularity.models.