readable AI benchmarks (simplified)

0 models · 0 points
filters graph + shortlist
Quality Score · answer quality plus multi-source capability · higher is better

loading live catalog · GET /api/recommend

Blended price · USD per 1M tokens · log scale · lower is betterBlended price · lower is better
most attractive Pareto frontierPoints use exact values. Labels move · data points do not. Only exact duplicates grouped.

Shortlist

click a model for full overview

#Model

loading ranking · GET /api/recommend

Value Score (v2) always uses blendedPrice, not selected graph cost, and requires a positive price plus Incorrect. Quality Score (v4) gives answerQuality a pre-normalization weight of 0.50; available weights always normalize. Intelligence uses 85% AA Intelligence plus 15% Epoch rank only on an exact family+release match. Coding uses 70% normalized AA Coding plus 30% CursorBench rank only for one exact canonical result. Agentic uses 70% normalized AA Agentic plus 30% site DeepSWE operational score only on an exact model+effort match; that score combines where pass@1 ranks with where fewer mean agent steps / pass@1 ranks. Missing cross-checks leave their normalized AA domain unchanged. AA indices are combined without source-overlap correction. Both are this site’s decision scores, not Artificial Analysis benchmarks.

ranked pick: GET /api/recommend · graph is not the ranking · API