About
this site is a benchmark aggregator. published scores, prices, and boards land in one catalog so you can compare models without hopping leaderboards.
I got tired of the constant juggling of different benchmark sites, so I decided to make my own graph of price vs quality and pick what I actually want to run.
Sources
- Artificial Analysis supplies broad model roster, token prices, Intelligence / Coding / Agentic indices, and AA-Omniscience correct / incorrect / abstain rates. See the Artificial Analysis methodology and language leaderboard.
- Models.dev enriches matched language, image, and video models with input and output modalities and supporting catalog metadata. It does not control benchmark scores, ranking, reasoning-effort variants, or comparison pricing; unmatched models remain valid. See the Models.dev catalog and source repository.
- Epoch AI General ECI is an independent graph/API cross-check on exact family+release matches. It does not enter Language Quality Score or fill missing AA Intelligence. See the Epoch methodology, included benchmarks, and ECI CSV.
- DeepSWE supplies independent software-engineering benchmark results on exact model and effort matches. Its pass@1, agent steps, cost, tokens, and duration remain separate graph/API dimensions; they do not enter Quality Score. See the official v1.1 data.
- CursorBench (source-reported version) provides an independent coding-agent graph/API result on an exact canonical match. Duplicate or missing matches remain absent. It does not enter Quality Score.
- Independent benchmark dimensions include DeepSWE, FrontierCode, CursorBench, Terminal-Bench, BullshitBench v2, and LMArena WebDev. Native values remain available on Graph and through
/api/benchmarks; sparse boards are not forced into Quality. - Provider plans use official pricing pages from OpenAI, Anthropic, Google, Xiaomi, Kimi, Z AI, MiniMax, and SpaceXAI.
- This site documents Quality Score, Value Score, and Factual reliability in its methodology.
- Image and Video comparison currently uses Artificial Analysis image and video quality and pricing data. Exact LM Arena media matches are shown separately as Additional Benchmarks from the official LM Arena leaderboard dataset. See Methodology for matching and scoring details.