readable AI benchmarks (simplified)
Quality Score · AA Intelligence and Omniscience outcomes · higher is better
loading live catalog · GET /api/recommend
Blended price · USD per 1M tokens · log scale · lower is betterBlended price · lower is better
most attractive Pareto frontierPoints use exact values. Labels move · data points do not. Only exact duplicates grouped.
#Model
loading ranking · GET /api/recommend
Quality Score (v6) combines AA Intelligence and AA-Omniscience Correct minus Incorrect outcomes at equal weight; both inputs are required. Value Score (v3) divides that Quality by a blended token-price penalty. Epoch, CursorBench, and DeepSWE remain separate graph/API dimensions and never fill missing Quality inputs. These are site decision scores, not published Artificial Analysis benchmarks. Methodology.
ranked pick: GET /api/recommend · graph is not the ranking · API