readable AI benchmarks (simplified)

0 models · 0 points
filters graph + shortlist
Quality Score · AA Intelligence and Omniscience outcomes · higher is better

loading live catalog · GET /api/recommend

Blended price · USD per 1M tokens · log scale · lower is betterBlended price · lower is better
most attractive Pareto frontierPoints use exact values. Labels move · data points do not. Only exact duplicates grouped.

Shortlist

click a model for full overview

#Model

loading ranking · GET /api/recommend

Quality Score (v6) combines AA Intelligence and AA-Omniscience Correct minus Incorrect outcomes at equal weight; both inputs are required. Value Score (v3) divides that Quality by a blended token-price penalty. Epoch, CursorBench, and DeepSWE remain separate graph/API dimensions and never fill missing Quality inputs. These are site decision scores, not published Artificial Analysis benchmarks. Methodology.

ranked pick: GET /api/recommend · graph is not the ranking · API

Support me! Patreon