readable AI benchmarks (simplified)
LongCat
LongCat 2.0
LongCat 2.0 has a Quality Score of 33.8, ranking 238th among 571 scored models. Per 1M tokens, pricing is $0.300 input, $1.20 output, and $0.0060 cached input. It accepts text input, outputs text, and has a 1M token context window. Its Reliability Score is 38.3, ranking 254th among 573 scored models, above average (total average: 36.5). Its Value Score is 28.8, around average compared with the total Value average of 31.4.
Quality Score33.8coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score28.8
Factual reliability38.3
Cache Discount98%
Model specification
- Reasoning
- reasoning
- Input modalities
- text
- Output modalities
- text
- Context window
- 1M tokens
- Weights
- open weights
Token prices USD per 1M tokens
- Input
- $0.300
- Output
- $1.20
- Cached Input
- $0.0060
- Cache write
- not available
- Cost per task
- $0.059
Capability
- Intelligence
- 19.1
- Coding
- not available
- Agentic
- not available
- Omniscience
- -23.4
- Correct
- 29.6%
- Blended price
- $0.184
Answer outcomes
Fully graded outcomes (Correct + Incorrect): 82.7%
Correct 29.6%Incorrect 53.0%Partial / not attempted 17.3%
Artificial Analysis benchmarks
- GDPval-AA v2
- 19%
- τ³-Banking
- 13%
- SciCode
- 36%
- Humanity’s Last Exam
- 34%
- GPQA Diamond
- 78%
- CritPt
- 3%
- AA-Omniscience
- 38%
- AA-LCR
- 65%
Similar models
- LongCat Flash LiteLongCatQuality Score 16.3Value Score not available
- Qwen3.6 35B A3B (Reasoning)AlibabaQuality Score 33.5Value Score not available
- GPT-5.4 nano (Xhigh)OpenAIQuality Score 33.6Value Score 28.6
- Grok 4.1 Fast (Reasoning)SpaceXAIQuality Score 33.2Value Score not available
- GPT-5.4 mini (Medium)OpenAIQuality Score 35.0Value Score 25.8
- GPT-5 mini (High)OpenAIQuality Score 33.6Value Score 27.4