readable AI benchmarks

52 points
Filters the graph and shortlist
Most attractive · 11 modelsGPT-5.6 Luna (max) · GPT-5.6 Luna (xhigh) · GPT-5.6 Luna (high) · DeepSeek V4 Pro (max) · Grok 4.5 (high) · DeepSeek V4 Pro (high) · DeepSeek V4 Flash (max) · Muse Spark 1.1 (xhigh) · Hy3 · DeepSeek V4 Flash (high) · GPT-5.6 Luna (medium)Balanced top 18% of enabled models by percentile rank on both axes · $1.4 or less · 41.0 or more
Price value · capability and reliability adjusted for price · higher is better
54
48
41
35
28
22
$0.04
$0.13
$0.38
$1.1
$3.2
$9.4
Blended price · USD per 1M tokens · log scale · lower is better
provider color + logo most attractive Pareto frontierPoints use exact values. Labels move; data points do not. Only exact duplicates are grouped.
Y-axis methodology

How each Y-axis metric is calculated

Site-calculated decision scores show their full formula. The Artificial Analysis indices are plotted exactly as published, so their benchmark composition stays aligned with the live source ↗.

Price value

This siteCurrent Y axis

A price-efficiency score that balances general capability and answer reliability, then applies a substantial affordability discount.

100 × [60% × clamp(Intelligence ÷ 65) + 40% × clamp(Correct × (1 − 0.35 × Incorrect))] × AffordabilityAffordability = 1 ÷ [1 + 0.45 × log₁₀(1 + 8 × blended price)]. Blended price is the 7:2:1 cache-hit, input, and output token mix. Intelligence, Correct, and blended price are required; a missing Incorrect value is treated as zero.

Quality value

This site

A quality-first score where answer outcomes dominate and affordability has only a small influence.

Weighted score: Answer quality 50% · Intelligence 20% · Coding 7.5% · Agentic 7.5% · Omniscience 5% · Evidence 7.5% · Affordability 2.5%Answer quality = clamp([0.60 × Correct + 0.40 × Correct ÷ (Correct + Incorrect)] × [1 − 0.75 × max(Incorrect − Correct, 0)]). Intelligence, Coding, and Agentic are normalized by 65, 80, and 60; Omniscience by (index + 50) ÷ 100. Correct and Incorrect are required. Unavailable optional components are removed and the remaining weights are renormalized.

Evidence score

Cross-source

A coverage-aware synthesis of independent leaderboards after each source is converted to an average-rank percentile, where the best result is 100.

Evidence = (General percentile + Software percentile + Truthfulness percentile + Preference percentile) ÷ 4General uses Artificial Analysis; Software averages DeepSWE, FrontierCode, CursorBench, and Terminal-Bench slots; Truthfulness uses BullshitBench; Preference uses LMArena WebDev. Missing slots contribute a neutral 50. A score is published only with at least four exact source matches across at least three families.

Correct answers

Artificial Analysis

The share of AA-Omniscience evaluation questions the model answered correctly, displayed as a percentage.

Correct answers = 100 × AA-Omniscience accuracy shareIncorrect answers and abstentions remain in the total question count but do not enter the correct-answer numerator. This site uses the published accuracy share without additional weighting.

Intelligence Index

Artificial Analysis

Artificial Analysis’s composite general-capability index across its current Intelligence Index evaluation suite.

Plotted value = published Artificial Analysis Intelligence IndexThe site does not recalculate, rescale, or mix this axis with price. Artificial Analysis owns the benchmark composition and publishes the index value used here.

Omniscience Index

Artificial Analysis

A net-correctness index that rewards correct answers and subtracts incorrect answers; abstentions contribute zero.

Omniscience Index = 100 × (Correct share − Incorrect share)The published Artificial Analysis index is plotted unchanged. The same relationship lets the site recover the incorrect-answer share when only Correct and the index are exposed.

Coding Index

Artificial Analysis

Artificial Analysis’s composite coding-capability index across its current coding evaluation suite.

Plotted value = published Artificial Analysis Coding IndexThe site applies no extra benchmark weighting or price adjustment to this axis. The source’s published numeric value is used unchanged.

Agentic Index

Artificial Analysis

Artificial Analysis’s composite measure of agentic and tool-using task performance.

Plotted value = published Artificial Analysis Agentic IndexThe site applies no extra benchmark weighting or price adjustment to this axis. The source’s published numeric value is used unchanged.

All answer shares are stored from 0 to 1 and displayed as percentages. “clamp” means a value is limited to the 0–1 range; a zero-answer precision is defined as zero.

Shortlist

Click a model to open its full overview.

#Model

Price value is the original price-efficiency score. Quality value is separate: 50% answer outcomes, 47.5% Intelligence Index, breakdowns and cross-source Evidence, and 2.5% affordability; its missing weight is redistributed when a component is unavailable. Models with more incorrect than correct answers receive an additional penalty. Both are this site’s decision scores, not Artificial Analysis benchmarks.