readable AI benchmarks

Meta

Llama 3.3 Instruct 70B

non-reasoning · open weights · Dec 6, 2024

Quality Score12.6
Value Score10.4
Reliability22.9
Cache Discount0%

Model specification

Reasoning
non-reasoning
Input modalities
text
Output modalities
text
Context window
128k tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.655
Output
$0.720
Cached Input
$0.655
Cached Output
not available
Cost per task
not available

Capability

Intelligence
9.3
Coding
11.9
Agentic
not available
Omniscience
-54.2
Correct
18.9%
Blended price
$0.662

Answer outcomes

Correct 18.9%Incorrect 73.1%Abstained 7.9%

Artificial Analysis benchmarks

GDPval-AA v2
0%
Terminal-Bench v2.1
5%
SciCode
26%
Humanity’s Last Exam
4%
GPQA Diamond
50%
CritPt
0%
AA-Omniscience
23%
AA-LCR
16%

Similar models