readable AI benchmarks

Alibaba

Qwen3.5 4B (Non-reasoning)

non-reasoning · open weights · Mar 2, 2026

Quality Score12.1
Value Score17.1
Reliability12.6
Cache Discountnot available

Model specification

Reasoning
non-reasoning
Input modalities
text, image, video
Output modalities
text
Context window
262,144 tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.030
Output
$0.150
Cached Input
not available
Cached Output
not available
Cost per task
not available

Capability

Intelligence
16.1
Coding
20.3
Agentic
not available
Omniscience
-74.7
Correct
11.2%
Blended price
$0.042

Answer outcomes

Correct 11.2%Incorrect 85.9%Abstained 2.8%

Artificial Analysis benchmarks

τ³-Banking
4%
Terminal-Bench v2.1
21%
SciCode
18%
Humanity’s Last Exam
8%
GPQA Diamond
71%
CritPt
0%
AA-Omniscience
13%
AA-LCR
34%

Similar models