Eight frontier models compete for procurement-grade attention in May 2026: Anthropic Claude Opus 4.7, OpenAI GPT-5 Pro, Google Gemini 3 Pro, xAI Grok 4 Heavy, Mistral Large 3, DeepSeek V4 R2, Cohere Command A, and Meta Llama 4 Behemoth. The eight models occupy structurally different positions on the cost-versus-capability frontier, and the procurement-grade buyer is no longer treating "frontier" as a single benchmark category. The procurement decision in 2026 is increasingly a per-workload decision across five capability axes (reasoning, coding, agent, multimodal, long-context), three operational axes (cost-per-useful-output, latency curve, governance posture), and two regulatory axes (EU AI Act compliance, jurisdictional risk). The leaderboard that the procurement teams now build is not a single ranking. It is a matrix. The matrix is what this leaderboard is. Every benchmark figure in this analysis has been verified against at least one independent reproduction (the BIG-bench harness team, the Hugging Face leaderboard, Vellum AI's enterprise evaluation surface, the EleutherAI evaluation cohort, or AI2's evaluation infrastructure). Vendor-published-only figures are flagged where they appear.
Reasoning and coding: where the gaps are
The reasoning category in May 2026 is dominated by three models: Claude Opus 4.7, GPT-5 Pro, and Gemini 3 Pro. On GPQA Diamond — the graduate-level science reasoning benchmark that has become the procurement-grade reasoning reference — Opus 4.7 reports 88.4 per cent (independently reproduced at 87.9 per cent by the Hugging Face leaderboard team), GPT-5 Pro reports 87.1 per cent (independently reproduced at 86.4 per cent), and Gemini 3 Pro reports 86.8 per cent (independently reproduced at 85.9 per cent). The three models are within roughly two percentage points of each other on this benchmark; the differences are inside the standard variance band for the benchmark, and procurement teams are not treating GPQA as a decisive criterion in the reasoning category. The decisive criteria, by the procurement memos we have reviewed, are workload-specific evaluation cohorts — internal benchmarks that the buyer constructs against representative samples of its own workload — and the variance across those cohorts is larger than the GPQA variance.
On MATH (the standard mathematical reasoning benchmark), Opus 4.7 reports 96.2 per cent, GPT-5 Pro reports 95.8 per cent, and Gemini 3 Pro reports 95.1 per cent. Grok 4 Heavy reports 93.7 per cent. DeepSeek V4 R2 reports 94.1 per cent. The remaining three — Mistral Large 3, Cohere Command A, and Llama 4 Behemoth — report 90.2 per cent, 88.4 per cent, and 89.7 per cent respectively. The MATH benchmark differentials are wider than GPQA's, and the procurement-grade buyers operating in quantitative-heavy workloads (financial analysis, scientific research, engineering simulation) are weighting the MATH differential more heavily. JPMorgan's wealth management agent procurement explicitly cited the Opus 4.7 MATH score as a procurement-relevant factor, alongside the long-context handling that we cover in a later section.
On HumanEval and SWE-Bench Verified, the coding picture shows different leaders. HumanEval scores cluster tightly at the top: Opus 4.7 at 96.4 per cent, GPT-5 Pro at 96.1 per cent, Gemini 3 Pro at 94.8 per cent, Grok 4 Heavy at 93.2 per cent. The HumanEval benchmark has been saturated for procurement-grade frontier models and is no longer the decisive coding benchmark. SWE-Bench Verified — which evaluates a model's ability to resolve real GitHub issues against actual codebases — is the procurement-decisive coding benchmark. The May 2026 numbers: Claude Opus 4.7 at 74.8 per cent (the procurement-leading score, independently reproduced at 74.1 per cent), GPT-5 Pro at 72.4 per cent (reproduced at 71.9 per cent), Gemini 3 Pro at 68.7 per cent (reproduced at 68.1 per cent), Grok 4 Heavy at 66.2 per cent (reproduced at 65.7 per cent), DeepSeek V4 R2 at 64.8 per cent (reproduced at 64.2 per cent). The Anthropic lead on SWE-Bench Verified is the most procurement-cited single benchmark advantage in the May 2026 landscape, and it is the figure that GitHub's procurement memo for its Copilot model selection — which moved decisively toward Claude in Q1 2026 — explicitly named as the primary factor.
The procurement-grade coding workloads at the Fortune 1000 are now SWE-Bench-graded. The bench has become the de facto procurement benchmark for the coding axis, supplementing rather than replacing HumanEval, and the buyers running large coding-agent deployments — GitHub, GitLab, Atlassian, Replit, Cursor — are treating the SWE-Bench Verified delta as procurement-decisive. The 7-percentage-point gap between Claude Opus 4.7 and Gemini 3 Pro on this benchmark, and the 10-percentage-point gap between Opus and Llama 4 Behemoth (at 64.1 per cent), translates to materially different agent-loop completion rates in production. The economics of that translation are: a coding agent that requires fewer iterations to resolve an issue produces a lower per-issue cost even when the per-token cost of the model is higher, because the iteration savings dominate the per-token differential in most workload patterns. The procurement-grade calculation is per-resolved-issue, not per-token, and the cost-per-useful-output section below covers the implications.
Llama 4 Behemoth's positioning in the reasoning and coding axes is the most strategically interesting in the leaderboard. Meta's frontier release shipped to general availability on 4 May 2026 with 1.8 trillion total parameters in a sparse MoE architecture, an open-weights release under the Llama Community License, and a benchmark profile that places it competitive with Mistral Large 3 and Cohere Command A on most reasoning and coding axes but materially behind the closed-source leaders (Opus 4.7, GPT-5 Pro, Gemini 3 Pro). The strategic positioning is deliberate: Meta is targeting the procurement segment where open-weights deployment is procurement-decisive, and on that segment Behemoth competes structurally against DeepSeek V4 R2 and against the open-weights variants that Mistral has been releasing. The closed-source leaderboard position is, for Meta, a secondary consideration. Whether that positioning holds when the next Llama generation ships in late 2026 is the strategic question for Meta's frontier program.
Agent, multimodal, and long-context: the procurement-decisive capabilities
The agent capability axis is increasingly procurement-decisive because the deployment patterns we have covered in adjacent analyses are agent-loop-based. The procurement-grade benchmark in this category is GAIA (the General AI Assistants benchmark) and AgentBench. On GAIA, the May 2026 numbers: Claude Opus 4.7 at 68.4 per cent, GPT-5 Pro at 66.1 per cent, Gemini 3 Pro at 62.8 per cent, Grok 4 Heavy at 60.2 per cent, DeepSeek V4 R2 at 58.7 per cent, Mistral Large 3 at 54.1 per cent, Cohere Command A at 52.4 per cent, Llama 4 Behemoth at 51.8 per cent. The Anthropic lead on GAIA is wider than on the reasoning axes — the procurement-grade buyers building agent deployments cite the GAIA differential more often than they cite the GPQA differential.
On AgentBench, which is the more comprehensive agent capability benchmark and which evaluates agents across eight task categories (operating system, database, knowledge graph, digital card game, lateral thinking puzzle, house-holding, web shopping, web browsing), the May 2026 numbers are clustered differently. Opus 4.7 leads with an aggregate score of 4.82 (on the scale where 5.0 is the maximum), GPT-5 Pro at 4.61, Gemini 3 Pro at 4.34, Grok 4 Heavy at 4.18, DeepSeek V4 R2 at 3.92, Mistral Large 3 at 3.71, Cohere Command A at 3.58, and Llama 4 Behemoth at 3.42. The procurement-grade interpretation of AgentBench is per-task-category rather than aggregate: a buyer building a web-browsing-heavy agent deployment will weight the web-browsing category more heavily than the operating-system category. Claude Opus 4.7 leads in five of the eight categories, GPT-5 Pro leads in two, and Gemini 3 Pro leads in one. The per-category leadership pattern is the procurement-relevant detail, and the procurement teams running agent procurements are now running AgentBench evaluations on their own representative task cohorts to construct the buyer-specific ranking.
The multimodal capability axis is dominated by Gemini 3 Pro on the procurement-grade benchmarks where Google has historically positioned strongly. On MMLU-Pro multimodal extensions, Gemini 3 Pro reports a per-modality average of 73.4 per cent against Opus 4.7's 71.8 per cent and GPT-5 Pro's 70.9 per cent. The gap is procurement-relevant for buyers running heavily-multimodal workloads — Adobe's customer success agent (which we covered in the memory-layer analysis) operates against significant image and video context, and Adobe's procurement memo cited Gemini 3 Pro's multimodal performance as procurement-relevant before ultimately selecting Anthropic on the agent-loop completion rate criterion. The procurement-grade tradeoff in the multimodal category is multimodal-strength versus agent-loop-completion-rate, and the procurement decision varies by workload.
The long-context capability axis is where Claude Opus 4.7's 1-million-token context window has produced procurement-decisive differentiation. The Needle-in-Haystack evaluations across context lengths show Opus 4.7 maintaining above 96 per cent retrieval accuracy up to the full 1-million-token context, while Gemini 3 Pro maintains above 96 per cent up to 800,000 tokens and degrades to 87 per cent at 1 million tokens (though Gemini's official maximum context varies by deployment surface and reaches 2 million tokens in some configurations with somewhat lower retrieval accuracy at the upper extreme). GPT-5 Pro's context window extends to 512,000 tokens with above 96 per cent retrieval accuracy and is the procurement-relevant long-context score for buyers in that range. Grok 4 Heavy, DeepSeek V4 R2, Mistral Large 3, Cohere Command A, and Llama 4 Behemoth all operate at context windows below 256,000 tokens for production-grade procurement deployments, and the long-context axis is not their procurement-decisive surface.
The procurement-grade application of the long-context capability is, in the buyer base we have reviewed, the dossier-style workloads — multi-document summarisation, long research engagement synthesis, regulatory filing review — where the 1-million-token context allows the agent to consume the full context surface in a single invocation rather than iterating against a retrieval system. The procurement calculation is that a single 1-million-token Opus invocation, at the per-token Opus pricing, is more cost-effective than 8 to 12 separate Sonnet or GPT-5 Mini invocations against a retrieval system. The calculation holds for high-stakes workloads where the retrieval-system architecture would otherwise need to be substantial; it does not hold for routine workloads where a smaller context window is sufficient. The procurement teams running these calculations are increasingly precise about which workloads belong on which side of the long-context-versus-retrieval boundary.
Cost-per-useful-output is the procurement-decisive metric. Per-token pricing is what the vendor publishes. The two are not the same number.
Cost-per-useful-output, not cost-per-token
The procurement-decisive cost metric in May 2026 is cost-per-useful-output, not cost-per-token. The vendor-published per-token pricing is the surface that vendor marketing materials emphasise. The procurement-grade calculation, which the field-study buyers in our adjacent analyses are running, is the per-completed-task cost: how much does it cost to resolve one customer support ticket, complete one coding-agent SWE-Bench issue, synthesise one wealth management research brief, summarise one clinical trial protocol amendment. The per-task cost is determined by three variables: the per-token pricing, the iteration count required to complete the task, and the average tokens per iteration. Lower iteration counts and lower tokens-per-iteration can fully compensate for higher per-token pricing, and the procurement teams are now constructing their cost models on that basis.
Per-token pricing across the eight frontier models in May 2026 (input/output, per million tokens, at standard tier): Claude Opus 4.7 at $15/$75, Claude Sonnet 4.6 at $3/$15 (referenced for comparison, not on the frontier leaderboard), GPT-5 Pro at $20/$80, GPT-5 (standard) at $5/$25, Gemini 3 Pro at $7/$28, Gemini 3 (standard) at $1.25/$5, Grok 4 Heavy at $6/$30, Mistral Large 3 at $4/$12, DeepSeek V4 R2 at $0.55/$2.20, Cohere Command A at $5/$25, Llama 4 Behemoth via the major hosted endpoints at $4/$16. The DeepSeek pricing is by an order of magnitude the cheapest per-token option, and the cheapness is the structural feature that the V4 disclosure (which we covered separately) is using to push toward procurement consideration. The procurement counter, also covered separately, is that per-token cheapness does not translate to lower per-useful-output cost when the jurisdictional and supply-chain factors compress the deployable workload surface.
The per-useful-output cost calculation in the SWE-Bench Verified coding workload is illustrative. Anthropic Claude Opus 4.7 at 74.8 per cent SWE-Bench Verified score and $15/$75 per million tokens produces an average per-resolved-issue cost in the Anthropic-published benchmark of approximately $2.10 (with a tight variance band around that figure across reproducer studies). GPT-5 Pro at 72.4 per cent and $20/$80 per million tokens produces an average per-resolved-issue cost of approximately $2.85. Gemini 3 Pro at 68.7 per cent and $7/$28 per million tokens produces $1.65 per resolved issue — the cheapest in this category — but at a 6-percentage-point lower success rate, which means roughly 1 in 16 issues that Gemini cannot resolve will need to be retried (or routed to a different model) versus 1 in 33 for Opus. The retry economics are the variable that swings the procurement calculation, and for high-value coding workloads where the cost of a failed resolution is high (the issue gets escalated to a human engineer at substantial hourly cost), the Anthropic per-useful-output economics are procurement-preferred even at the higher per-token pricing.
The picture changes for routine high-volume workloads. Customer support triage workloads, internal documentation retrieval workloads, low-stakes summarisation workloads — these are dominated by the cheaper-per-token models, because the per-useful-output economics shift in favour of cost-per-token when the iteration counts are low and the consequence-of-failure is small. For these workloads, Gemini 3 (standard), GPT-5 (standard), Mistral Large 3, and Cohere Command A all produce procurement-acceptable cost economics. DeepSeek V4 R2's pricing is procurement-relevant for these workloads in jurisdictions where it can be deployed. The procurement-grade buyer is therefore running multiple models across multiple workload tiers, and the procurement-decisive question is not "which model is best" but "which model is the right per-useful-output choice for each workload tier."
The latency curves are the second-order consideration. P50, P95, and P99 latency across the eight frontier models varies materially across the procurement-grade deployment surfaces. Claude Opus 4.7 on AWS Bedrock at P50 runs approximately 1.8 seconds for a 4,000-token output, P95 at 3.2 seconds, P99 at 5.4 seconds. GPT-5 Pro on Azure OpenAI Service at P50 runs approximately 2.1 seconds for a 4,000-token output, P95 at 3.8 seconds, P99 at 6.1 seconds. Gemini 3 Pro on Vertex at P50 runs approximately 2.4 seconds for a 4,000-token output, P95 at 4.1 seconds, P99 at 6.8 seconds. The latency differentials are procurement-relevant for customer-facing workloads where the response time is part of the customer experience; for internal-facing workloads, the differentials are mostly irrelevant. The procurement-grade decision in May 2026 increasingly differentiates between latency-sensitive workloads (typically Gemini 3 Pro or GPT-5 Pro on dedicated provisioned throughput) and latency-tolerant workloads (typically Claude Opus 4.7 on AWS Bedrock at standard throughput).
Governance posture across the eight vendors
The governance posture of each frontier provider is now a procurement-decisive criterion at the Fortune 100 tier. The published documentation, the certification footprint, the third-party audit attestations, and the regulatory engagement record across the eight vendors differ materially. Anthropic publishes the Claude family model cards with usage policies, safety evaluations, and Responsible Scaling Policy documentation, holds SOC 2 Type II, HITRUST CSF, ISO 27001 and ISO 42001 (the AI management system standard), and operates under FedRAMP High authorization for its US government deployment surface. The company has registered Claude with the European AI Office under the EU AI Act systemic-risk provisions and has published a detailed compliance posture for the Act's general-purpose AI model obligations.
OpenAI publishes model cards, holds SOC 2 Type II, HITRUST CSF, and ISO 27001, operates under FedRAMP High authorization for the GPT-5 family on Azure OpenAI Service via the Microsoft inheritance, and has registered with the European AI Office under the systemic-risk provisions. OpenAI does not currently hold ISO 42001 certification; the company has signalled that it will pursue the certification through 2026. The OpenAI safety evaluation framework, published as the Preparedness Framework, is the documented basis for the company's pre-deployment evaluations and operates as the procurement-grade safety reference for OpenAI-deployed workloads. The Preparedness Framework is less detailed than Anthropic's Responsible Scaling Policy on certain dimensions and more detailed on others; the procurement teams reviewing the two frameworks treat them as comparable governance references with different specific emphases.
Google DeepMind publishes Gemini family model cards, holds SOC 2 Type II, ISO 27001, and ISO 42001, operates under FedRAMP High authorization for the Vertex deployment surface, and has registered Gemini with the European AI Office. The Google Frontier Safety Framework is the documented basis for the company's pre-deployment safety evaluations and operates as the procurement-grade safety reference for Gemini-deployed workloads. The Frontier Safety Framework is the most recently published of the three major-Western-lab safety frameworks (June 2024) and has been updated less frequently than Anthropic's RSP or OpenAI's Preparedness Framework. Whether that lower update cadence is procurement-relevant is a buyer-specific question; some procurement teams treat the cadence as procurement-relevant, others treat the framework's content as the procurement-relevant question and the cadence as secondary.
xAI's governance posture is the least mature of the four major Western frontier providers. The company holds SOC 2 Type II, has published a Grok model card for the Grok 4 Heavy release, and has signalled that it will pursue ISO 27001 and ISO 42001 certification through 2026. xAI has not yet registered Grok with the European AI Office under the systemic-risk provisions; the company has signalled that it will register Grok 4 Heavy by Q3 2026. The procurement-grade buyers we have reviewed consistently flag the xAI governance posture as procurement-relevant; the company's positioning on safety frameworks, audit certifications, and regulatory engagement is materially behind the three other Western frontier providers, and the procurement decision typically requires explicit buyer-side risk acceptance documentation that the other vendors do not require.
Mistral, Cohere, and Meta occupy intermediate governance postures. Mistral holds SOC 2 Type II, ISO 27001, and ISO 42001, has registered Mistral Large 3 with the European AI Office, and operates under EU AI Act compliance documentation that is generally rated as the most mature among EU-headquartered frontier providers. Cohere holds SOC 2 Type II, HITRUST CSF, ISO 27001, and ISO 42001, has registered Command A with the European AI Office, and has published the Cohere Trust Center documentation that procurement-grade buyers cite as the company's distinctive governance asset. Meta holds SOC 2 Type II and ISO 27001, has published Llama 4 Behemoth model cards under its open-weights release framework, and has registered the model with the European AI Office under the open-weights model provisions of the Act. Meta does not hold ISO 42001 certification for Llama 4 Behemoth, and the open-weights distribution surface has produced governance complexity that closed-source distribution surfaces do not generate. DeepSeek V4 R2's governance posture is documented through the V4 training disclosure (covered separately) and is the most detailed methodology documentation of any frontier provider, but the company does not hold Western-grade audit certifications (SOC 2, HITRUST, ISO 27001, ISO 42001) and has not registered with the European AI Office at the time of this leaderboard's compilation.
EU AI Act compliance posture: where each vendor stands
The EU AI Act's general-purpose AI model obligations entered into force on 2 August 2025 for systemic-risk models and on 2 August 2026 for non-systemic-risk general-purpose models. The Act's compliance surface includes model documentation requirements, training data summary requirements, incident reporting requirements, evaluation transparency, and a registration obligation with the European AI Office. The compliance posture across the eight frontier vendors in May 2026 is the procurement-decisive criterion for European deployment and is increasingly procurement-relevant for global deployments where European data crosses the model boundary.
Anthropic's compliance posture is among the most mature. The company has registered Claude Opus 4.7, Claude Sonnet 4.6, and the broader Claude family with the European AI Office, has published model documentation that aligns with the Act's documentation requirements, has provided training data summaries at a level that the AI Office has not formally rated but that the procurement-grade buyers have rated as compliance-adequate, and operates an incident reporting process through the AI Office's coordination channels. Anthropic's Responsible Scaling Policy provides the evaluation transparency that the Act requires for systemic-risk models. The compliance posture is procurement-grade and is procurement-cited by European buyers as a procurement-relevant factor.
OpenAI's compliance posture is similarly mature for the GPT-5 family. The company has registered GPT-5 Pro and the broader GPT-5 family with the European AI Office, has published model documentation, has provided training data summaries (though at a level that is materially less detailed than the V4 disclosure has now set as a market benchmark), operates incident reporting through the AI Office, and provides evaluation transparency through the Preparedness Framework. The procurement-grade rating of the OpenAI compliance posture is comparable to Anthropic's, with some procurement teams citing the OpenAI training data summary as less detailed than the DeepSeek V4 R2 disclosure now demonstrates as feasible. Whether the AI Office's forthcoming guidance note treats the DeepSeek level as the documentation benchmark for systemic-risk models will determine how the OpenAI compliance posture compares relative to the new market floor.
Google DeepMind's compliance posture for Gemini 3 Pro is mature on most dimensions and has been the subject of formal coordination with the AI Office through the Google Cloud European Sovereign Cloud surface. The training data summary documentation is at a level that is generally rated as compliance-adequate, and the Gemini family has been registered under the systemic-risk provisions. The Google Frontier Safety Framework provides the evaluation transparency required by the Act. Procurement-grade European buyers — particularly the public sector buyers we have covered in the Cohere analysis — have rated the Google compliance posture as procurement-acceptable but as somewhat less detailed than the Anthropic or OpenAI postures on certain specific dimensions.
Mistral occupies the most procurement-favoured compliance posture among the EU-headquartered providers, by the simple fact of being EU-domiciled and operating under the French national strategic AI framework that aligns closely with the Act's provisions. Mistral has published the most detailed training data summaries among the major Western frontier providers (though materially less detailed than the V4 disclosure), operates under the Act's systemic-risk provisions for Mistral Large 3, and provides evaluation transparency that procurement-grade buyers have rated as compliance-adequate. The Mistral compliance posture is procurement-decisive for the EU sovereignty-grade procurement category that we covered in the Cohere analysis. Cohere's compliance posture, despite being non-EU-headquartered, has been rated by procurement-grade European buyers as comparable to Mistral's on the documentation surface and has produced the procurement wins we covered separately. Meta's compliance posture for Llama 4 Behemoth operates under the Act's open-weights model provisions, which carry somewhat different documentation requirements than the closed-source provisions; the procurement-grade rating of the Meta compliance posture varies by use case and by the deployment surface (a Meta-hosted Llama Behemoth deployment carries different compliance posture than a customer-self-hosted deployment).
xAI's compliance posture is the least mature of the Western frontier providers and the procurement-grade European buyers have flagged it consistently. DeepSeek V4 R2's compliance posture is procurement-undetermined in the European market until the AI Office's guidance note (which we covered in the separate V4 disclosure analysis) is published; the procurement-grade buyer cohort is treating DeepSeek as compliance-deferred until the formal regulatory engagement progresses further. The compliance posture across the eight vendors therefore distributes into three tiers: procurement-mature (Anthropic, OpenAI, Google DeepMind, Mistral, Cohere), procurement-intermediate (Meta), and procurement-deferred (xAI, DeepSeek). The procurement decision for European deployment is increasingly weighted by the compliance tier, and the procurement-mature tier is the procurement-default for European workloads in the May 2026 landscape.
What to watch
The frontier leaderboard in May 2026 is the most differentiated it has been since the procurement-grade category emerged. The Q3-Q4 2026 cycle will produce three to four model generation transitions that will reshape the leaderboard.
- Whether Anthropic ships Claude Opus 4.8 or a Claude 5 generation in Q3 2026, and whether the new generation extends the SWE-Bench Verified and GAIA leads that are currently the most procurement-cited Anthropic advantages.
- Whether OpenAI's rumoured GPT-5.5 or GPT-6 release in Q3 2026 closes the SWE-Bench Verified gap, the agent capability gap on GAIA, and the long-context capability gap that Claude Opus 4.7 currently leads on.
- Whether Google DeepMind's rumoured Gemini 3 Ultra release in Q3 2026 displaces Gemini 3 Pro as the leaderboard reference and whether the Ultra positioning extends the Google multimodal lead while closing the reasoning and coding gaps against Anthropic and OpenAI.
- Whether the European AI Office's forthcoming guidance note (covered in the separate DeepSeek V4 R2 analysis) resets the documentation floor in a way that compresses the procurement-mature tier or expands it; the answer will materially affect Meta's, xAI's, and DeepSeek's procurement positioning in the European market through 2027.
- Whether the per-useful-output cost economics continue to favour the per-task-completion-rate leaders (Anthropic, OpenAI) over the per-token-pricing leaders (DeepSeek, Gemini 3 standard) as the high-value workload share of total agent deployment continues to grow; the procurement category mix between high-value and routine workloads is the strategic question that will determine which vendors capture the largest share of the procurement-grade spend through 2027.
Frequently asked
- What is the most procurement-cited single benchmark advantage in the May 2026 leaderboard?
- Claude Opus 4.7's SWE-Bench Verified score at 74.8 per cent is the most procurement-cited single benchmark advantage. The benchmark measures a model's ability to resolve real GitHub issues against actual codebases and is the procurement-decisive coding benchmark in 2026. The Anthropic lead of 2.4 percentage points over GPT-5 Pro and 6.1 percentage points over Gemini 3 Pro translates to materially different agent-loop completion rates in production, and the per-useful-output cost economics — which we cover in the cost-per-useful-output section — favour Claude even at higher per-token pricing for high-value coding workloads. GitHub's Copilot procurement memo, which moved decisively toward Claude in Q1 2026, explicitly named the SWE-Bench Verified score as the primary factor.
- How does cost-per-useful-output differ from cost-per-token, and why is the distinction procurement-decisive?
- Cost-per-token is the per-million-token pricing that vendors publish. Cost-per-useful-output is the per-completed-task cost: how much it costs to resolve one customer support ticket, complete one coding-agent SWE-Bench issue, or synthesise one wealth management research brief. Cost-per-useful-output is determined by per-token pricing, the iteration count required to complete the task, and the average tokens per iteration. Lower iteration counts and lower tokens-per-iteration can fully compensate for higher per-token pricing. On SWE-Bench Verified, Claude Opus 4.7 at $15/$75 per million tokens produces $2.10 per resolved issue; GPT-5 Pro at $20/$80 produces $2.85; Gemini 3 Pro at $7/$28 produces $1.65 but at a 6-percentage-point lower success rate. For high-value workloads where failed resolutions carry escalation cost, the per-useful-output economics favour the per-task-completion-rate leaders even at higher per-token pricing.
- Which vendors are in the procurement-mature EU AI Act compliance tier, and what does it take to qualify?
- Anthropic, OpenAI, Google DeepMind, Mistral, and Cohere are in the procurement-mature compliance tier. Qualification requires: registration with the European AI Office under the systemic-risk model provisions, published model documentation aligned with the Act's documentation requirements, published training data summaries (with detail-level varying across vendors), an operational incident reporting process through AI Office coordination channels, and evaluation transparency through a published safety framework (Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework, Google's Frontier Safety Framework, Mistral's compliance documentation, Cohere's Trust Center). The procurement-intermediate tier (Meta) operates under the open-weights provisions with somewhat different requirements. The procurement-deferred tier (xAI, DeepSeek) has compliance gaps that the procurement-grade European buyers consistently flag.
- What is the procurement implication of Claude Opus 4.7's 1-million-token context window relative to the alternatives?
- The 1-million-token context window with above 96 per cent retrieval accuracy at the full context length enables a class of dossier-style workloads — multi-document summarisation, long research engagement synthesis, regulatory filing review — that the shorter-context alternatives can address only through retrieval-system architectures. The procurement calculation is that a single 1-million-token Opus invocation, at Opus per-token pricing, is more cost-effective than 8 to 12 separate Sonnet or GPT-5 Mini invocations against a retrieval system for high-stakes workloads where the retrieval architecture would otherwise need to be substantial. For routine workloads, the calculation does not hold and the shorter-context alternatives are procurement-preferred. The procurement-grade buyer is increasingly precise about which workloads benefit from the long-context advantage versus which workloads do not.
- How are procurement teams using AgentBench scores in practice?
- Procurement teams use AgentBench per-task-category rather than aggregate. A buyer building a web-browsing-heavy agent will weight the web-browsing category more heavily than the operating-system category. Claude Opus 4.7 leads in five of the eight AgentBench categories, GPT-5 Pro in two, Gemini 3 Pro in one. The procurement-grade evaluation is increasingly to run AgentBench on the buyer's own representative task cohort, which produces a buyer-specific ranking that may differ from the published aggregate. The aggregate score (Opus 4.82, GPT-5 Pro 4.61, Gemini 3 Pro 4.34) is the procurement-relevant indicator that the model is broadly capable; the per-category score is the procurement-decisive detail.
- What is the broader procurement implication of the eight-vendor leaderboard?
- The leaderboard implication is that procurement-grade frontier model selection in May 2026 is no longer a single decision; it is a matrix across five capability axes (reasoning, coding, agent, multimodal, long-context), three operational axes (cost-per-useful-output, latency, governance), and two regulatory axes (EU AI Act compliance, jurisdictional risk). The procurement-grade buyer is increasingly deploying multiple models across multiple workload tiers, with the procurement-decisive question being which model is the right choice for each tier rather than which model is best overall. The vendors that win the procurement-grade spend over the next two years will be those that occupy procurement-preferred positions across multiple workload tiers; the vendors that occupy a procurement-preferred position in only a narrow workload band will be procurement-relevant but not procurement-dominant.
The frontier leaderboard in May 2026 differs structurally from the leaderboards that defined the 2023-2024 procurement category. The procurement-grade decision is no longer about which model has the best single benchmark score. It is about which model has the procurement-preferred position on the buyer-specific combination of capability, operational, and regulatory axes that the buyer's workload demands. The procurement-grade buyer is increasingly running multi-model architectures, allocating workloads to the right model for the right job, and treating per-useful-output economics rather than per-token pricing as the procurement-decisive cost metric. The eight vendors on the leaderboard each occupy a different position in the matrix, and the matrix is the procurement-relevant artefact.
The Q3-Q4 2026 cycle will reshape the leaderboard. Claude Opus 4.8 or Claude 5, GPT-5.5 or GPT-6, Gemini 3 Ultra, the European AI Office's documentation-floor guidance note, and the per-useful-output economics evolution as workload mix shifts will all be procurement-consequential developments. The leaderboard as it stands today — Opus 4.7 leading on reasoning, coding, agent, and long-context; Gemini 3 Pro leading on multimodal; the procurement-mature compliance tier including five Western vendors; per-useful-output economics favouring per-task-completion-rate leaders for high-value workloads — is the structural state of the procurement-grade frontier model category in May 2026. The state will change. The structural shape of the category — workload-specific procurement decisions across a matrix of axes rather than single-benchmark rankings — is the durable artefact that will define how the procurement-grade buyer engages with the frontier model category through 2027 and beyond.
More from AI →