readable AI benchmarks (simplified)

readable AI benchmarks

Meta

Llama 3.3 Instruct 70B

non-reasoning · open weights · Dec 6, 2024

Llama 3.3 Instruct 70B has a Quality Score of 17.4, ranking 460th among 571 scored models. Per 1M tokens, pricing is $0.710 input, $0.720 output, and $0.710 cached input. It has a 128k token context window. Its Reliability Score is 22.9, ranking 435th among 573 scored models, below average (total average: 36.5). Its Value Score is 12.7, low compared with the total Value average of 31.4.

Quality Score17.4coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Score12.7
Factual reliability22.9
Cache Discount0%

Model specification

Reasoning
non-reasoning
Input modalities
none
Output modalities
none
Context window
128k tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.710
Output
$0.720
Cached Input
$0.710
Cache write
not available
Cost per task
not available

Capability

Intelligence
7.7
Coding
not available
Agentic
not available
Omniscience
-54.2
Correct
18.9%
Blended price
$0.711

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 92.1%

Correct 18.9%Incorrect 73.1%Partial / not attempted 7.9%

Artificial Analysis benchmarks

GDPval-AA v2
0%
Humanity’s Last Exam
4%
GPQA Diamond
50%
CritPt
0%
AA-Omniscience
23%
AA-LCR
16%

Similar models

Support me! Patreon