readable AI benchmarks (simplified)

Overview

This page documents the json routes. The Graph on / is a selectable scatter, not the ranking. Its axes include site scores, pricing, capability, and independent benchmarks such as DeepSWE Pass@1 and Avg agent steps. Use GET /api/models for the full catalog with Quality values and GET /api/recommend for ranked selection. No API key.

  • /api/models models[] is the full catalog for bulk integrations
  • /api/recommend models[] is the requested ranking (default: Quality Score)
  • price[] is Value Score (quality per dollar). It is not cheapest Blended price.
  • cheap[] is lowest Blended price above 0
  • lookup: ?q=NAME (name / slug / provider)

Rankings

FieldTypeDescription
Quality Scorenumber 0-100requires finite Intelligence, Correct, Incorrect; Correct and Incorrect must be nonnegative and sum to at most 1; Quality Score = 100 × (0.5 × clamp01(Intelligence / 65) + 0.5 × clamp01((1 + Correct minus Incorrect) / 2)) Formula v6; requires finite AA Intelligence and valid AA-Omniscience Correct and Incorrect shares. It combines Intelligence / 65 and the net Correct minus Incorrect outcome at equal weight. Both required inputs must be present: absent data yields null, not an external fallback. Epoch General ECI, CursorBench, and DeepSWE remain separate benchmark graph/API dimensions and do not enter Quality or Value. Epoch includes DeepSWE in its benchmark basket, so those independent axes contain overlapping evidence. Quality Coverage is 1 only for eligible models; otherwise null. Correct + Incorrect counts fully graded outcomes; residual may include partial answers and non-attempts. Higher is better.
Value Scorenumber 0-100requires valid Quality Score and finite blendedPrice > 0; Value Score = Quality Score / (1 + 0.45 × log10(1 + 8 × blendedPrice)) Formula v3; always uses blendedPrice, not selected graph cost. Zero/local/unpublished prices are ineligible. higher is better.
Factual reliabilitynumber 0-100Factual reliability = clamp((correct minus incorrect + 1) / 2, 0, 1) × 100; requires finite correct and incorrect; 50 is neutral net outcome, not 50% correctness; residual may include partial answers and non-attempts Formula v1; Graph Y-axis only. higher is better.
Blended priceUSD / 1M tokens7:2:1 cache-hit · input · output. lower is better.
DeepSWE Avg agent stepsstepsSource mean_agent_steps: average mini-swe-agent steps across every scored attempt, including scored failures. lower is better as an efficiency dimension. Graph uses the exact attached model+effort value; missing or duplicate matches stay absent.
DeepSWE Mean cost per attemptUSD / attemptSource mean_cost_usd: mean benchmark cost in USD over every scored attempt under the source's cost basis. It is not Artificial Analysis cost per task and not token price (USD per 1M tokens). informational only; not a Quality Score or Value Score input.
DeepSWE Mean output tokens per attempttokens / attemptSource mean_output_tokens: mean output tokens over every scored attempt, including scored failures. informational only; not a Quality Score or Value Score input.
DeepSWE Mean duration per attemptsecondsSource mean_duration_seconds: mean wall-clock seconds per scored attempt. It varies with provider and host load and is informational only; not a Quality Score or Value Score input.
DeepSWE Mean cost per attemptUSD per attemptSource mean_cost_usd: mean benchmark cost in USD over every scored attempt under the source's cost basis. It is not Artificial Analysis cost per task and not token price (USD per 1M tokens). informational only; not a Quality Score, Value Score, or Reliability input. lower is not scored.
DeepSWE Mean output tokens per attempttokens per attemptSource mean_output_tokens: mean output tokens over every scored attempt, including scored failures. informational only; not a Quality Score, Value Score, or Reliability input.
DeepSWE Mean duration per attemptsecondsSource mean_duration_seconds: mean wall-clock seconds per scored attempt. It varies with provider and host load and is informational only; not a Quality Score, Value Score, or Reliability input.
Intelligence Indexnumberlive data Intelligence Index on the Graph and Shortlist.
Coding Indexnumberlive data Coding Index.
Agentic Indexnumberlive data Agentic Index.
Correct answersrate 0-1share of Omniscience questions answered correctly.
Incorrect answersrate 0-1share of Omniscience questions answered incorrectly.

cheap[] skips blendedPrice of 0 (local / unpublished rows). Value Score does too.

Authentication

None. GET, HEAD, OPTIONS. CORS *.

GET/api/models

Catalog endpoint for every explorer dataset. Omit dataset or use ?dataset=language for the existing language response, including canonical Quality Score and Quality Coverage. Use ?dataset=image or ?dataset=video for the Image and Video catalogs (Quality remains qualityElo), and ?dataset=all for all three named envelopes. Image and video rows may also include externalBenchmarks for exact LM Arena Score matches. That field is additive evidence and does not change qualityEloor default ranking. Envelope source / stale describe the primary Artificial Analysis media feed only. LM Arena live / cached / snapshot status lives in additive externalBenchmarkSources; a live AA catalog does not imply a live Arena fetch. It is a separate comparison group, not a ranking; use /api/recommend for ranked language selection.

curl -sS "https://models.deggo.fyi/api/models?dataset=image"
curl -sS "https://models.deggo.fyi/api/models?dataset=video"
curl -sS "https://models.deggo.fyi/api/models?dataset=all"

Language responses keep the existing source, updatedAt, models, and message envelope. Image and video responses add sourceUrl and stale for the primary catalog, plus externalBenchmarkSources for attached external boards. The all-dataset response returns those complete envelopes under datasets.language, datasets.image, and datasets.video. Exact language DeepSWE attachments expose deepSwe.meanAgentSteps, deepSwe.meanCostUsd, deepSwe.meanOutputTokens, and deepSwe.meanDurationSeconds (each number or null); the Graph reads that same raw steps value. Every language models[] row includes:

  • slug: canonical model slug
  • name: display name or null
  • provider: creator name
  • qualityValue: current Quality Score, number 0-100 or null when ineligible
  • qualityCoverage: 1 when required AA inputs are valid, otherwise null
  • attemptRate: Correct + Incorrect, number 0-1 or null
  • priceValue: current Value Score, number 0-100 or null

Existing catalog pricing, benchmark, modality, and capability fields remain unchanged. Missing Quality is returned as null, never 0, and no catalog rows are omitted.

Example Request

curl -sS "https://models.deggo.fyi/api/models"

GET/api/recommend

ranked pick for a constraint set. Default goal=quality, limit=10 (maximum 30). Use this for selecting models, not mirroring the full catalog; bulk catalog consumers should use /api/models. Default body also includes price[] and cheap[]. read is the first field after source.

Query

KeyTypeDefaultDescription
goalquality | price | cheap | coding | intelligencequalityquality: Quality Score. price: Value Score (not cheapest USD). cheap: lowest Blended price above 0. coding / intelligence: those indexes.
maxBlendedPricefinite numbernoneUSD / 1M Blended price ceiling.
minIntelligencefinite numbernoneminimum Intelligence Index.
minQualityValuefinite numbernoneminimum Quality Score.
reasoningtrue | falsenonekeep only reasoning or non-reasoning rows.
openWeightstrue | falsenoneopen or closed weights.
minContextWindowinteger tokensnoneminimum context window.
inputModalitycomma list text,image,video,speechnoneAND: model must include every listed input.
outputModalitycomma list text,image,video,speechnoneAND: model must include every listed output.
providerstring, max 64nonecase-insensitive provider substring.
qstring, max 64nonelookup by name / slug / provider. example: ?q=grok
limitinteger 1-3010row cap. 31 is 400.

Unknown keys return 400.

Response fields

FieldTypeDescription
sourcestringalways recommend
readstringfirst field after source. models[] is the requested ranking. price[] is Value Score, not cheapest USD. cheap[] is lowest Blended price. lookup: ?q=NAME. Full catalog with Quality values: GET /api/models.
goalstringapplied goal.
updatedAtISO datetimefrom the models feed, else now.
notestringcatalog provenance. HTML is a graph view, not the ranking. do not scrape /.
filtersobjectecho of applied query keys.
countintegerlength of models[].
modelsarrayrequested ranking.
pricearrayValue Score ranking. omitted when goal=price.
cheaparraylowest Blended price above 0. omitted when goal=cheap.

models[] row

FieldTypeDescription
slugstringcanonical slug.
namestringdisplay name.
providerstringcreator name.
reasoningboolean or nullreasoning row, if listed.
openWeightsboolean or nullopen weights, if listed.
contextWindowTokensinteger or nullcontext window.
inputModalitiesstring[]text / image / video / speech.
outputModalitiesstring[]text / image / video / speech.
blendedPricenumber or nullBlended price, USD / 1M.
intelligencenumber or nullIntelligence Index.
codingnumber or nullCoding Index.
agenticnumber or nullAgentic Index.
correctnumber or nullCorrect answers rate.
incorrectnumber or nullIncorrect answers rate.
qualityValuenumber or nullQuality Score 0-100.
qualityCoveragenumber 0-1 or null1 for valid AA Intelligence and AA-Omniscience outcomes; otherwise null.
attemptRatenumber 0-1 or nullCorrect + Incorrect.
priceValuenumber or nullValue Score 0-100.
whystringone short sentence from actual numbers.

Example Request

curl -sS "https://models.deggo.fyi/api/recommend"
curl -sS "https://models.deggo.fyi/api/recommend?q=grok"
curl -sS "https://models.deggo.fyi/api/recommend?goal=cheap"
from urllib.request import Request, urlopen

req = Request(
    "https://models.deggo.fyi/api/recommend",
    headers={"Accept": "application/json"},
)
with urlopen(req) as response:
    print(response.read().decode())

Example Response

{
  "source": "recommend",
  "read": "models[] is the requested ranking. default goal=quality. price[] is quality per dollar, not cheapest USD. cheap[] is lowest blendedPrice above 0. lookup one model: GET /api/recommend?q=NAME. full catalog with Quality values: GET /api/models. HTML is a graph view, not the ranking.",
  "goal": "quality",
  "updatedAt": "<from models feed, else null>",
  "catalogSource": "<live|cached|snapshot|null>",
  "catalogUpdatedAt": "<original models timestamp or null>",
  "catalogStale": true,
  "catalogMessage": "<models message or null>",
  "note": "catalog provenance. HTML is a graph view, not the ranking. do not scrape /.",
  "filters": { "goal": "quality", "limit": 10 },
  "count": 10,
  "models": [
    {
      "slug": "",
      "name": "",
      "provider": "",
      "reasoning": true,
      "openWeights": false,
      "contextWindowTokens": 0,
      "inputModalities": [],
      "outputModalities": [],
      "blendedPrice": 0,
      "intelligence": 0,
      "coding": 0,
      "agentic": 0,
      "correct": 0,
      "incorrect": 0,
      "qualityValue": 0,
      "qualityCoverage": 1,
      "attemptRate": 0,
      "priceValue": 0,
      "why": "short sentence from actual numbers"
    }
  ],
  "price": [],
  "cheap": []
}

Errors

StatusBodyDescription
400{ "error": "unknown query key: ..." }
{ "error": "invalid limit" }
{ "error": "query too long" }
unknown key, out of range, or query string over 512.
429{ "error": "rate limit" }Retry-After: 60. 30 / 60s per IP. 120 / 60s global.
503{ "error": "models unavailable" }catalog loader failed.
405{ "error": "method not allowed" }not GET, HEAD, or OPTIONS.

GET/api

Machine index. Fetch this first if query keys are unknown. Field use explains models[] vs price[] vs cheap[].

Example Request

curl -sS "https://models.deggo.fyi/api"

Other routes

  • /api/models full catalog the Graph uses, including canonical Quality Score and Quality Coverage. Image and video catalogs keep Quality on qualityElo and may attach exact LM Arena rows in externalBenchmarks. Primary feed status stays on source/stale; Arena status is externalBenchmarkSources. Use it for bulk/catalog integrations; cache response. It is not a ranking.
  • /api/benchmarks external boards: source/version/status metadata and native records. Each record includes canonicalSlug, matchConfidence, metric, rawValue, unit, and within-board rank when available. Missing results stay absent.
  • /api/subscriptions provider plans
  • /llms.txt crawler pointer

Limits

  • 30 / 60s per IP on /api and /api/recommend
  • 120 / 60s global
  • query string <= 512 characters
  • limit 1-30
  • no API key
  • GET / HEAD / OPTIONS only
Support me! Patreon