readable AI benchmarks (simplified)

readable AI benchmarks

Alibaba

Qwen3 VL 235B A22B Instruct

non-reasoning · open weights · Sep 23, 2025

Qwen3 VL 235B A22B Instruct has a Quality Score of 19.5, ranking 397th among 551 scored models. Per 1M tokens, pricing is $0.400 input and $1.60 output. It accepts text and image input, outputs text, and has a 262,144 token context window. Its Reliability Score is 23.7, ranking 402nd among 553 scored models, below average (total average: 35.6).

Quality Score19.5coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Scorenot available
Factual reliability23.7
Cache Discountnot available

Model specification

Reasoning
non-reasoning
Input modalities
text, image
Output modalities
text
Context window
262,144 tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.400
Output
$1.60
Cached Input
not available
Cache write
not available
Cost per task
not available

Capability

Intelligence
9.9
Coding
not available
Agentic
not available
Omniscience
-52.7
Correct
20.3%
Blended price
not available

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 93.3%

Correct 20.3%Incorrect 73.0%Partial / not attempted 6.7%

Artificial Analysis benchmarks

Humanity’s Last Exam
7%
GPQA Diamond
71%
CritPt
0%
AA-Omniscience
24%
AA-LCR
33%

Similar models

Support me! Patreon