readable AI benchmarks (simplified)

readable AI benchmarks

Meta

Llama 3.1 Instruct 8B

non-reasoning · open weights · Jul 23, 2024

Llama 3.1 Instruct 8B has a Quality Score of 22.6, ranking 333rd among 551 scored models. Per 1M tokens, pricing is $0.020 input and $0.050 output. It has a 128k token context window. Its Reliability Score is 34.6, ranking 264th among 553 scored models, below average (total average: 35.6).

Quality Score22.6coverage 100.0%: AA Intelligence and AA-Omniscience outcomes available
Value Scorenot available
Factual reliability34.6
Cache Discountnot available

Model specification

Reasoning
non-reasoning
Input modalities
none
Output modalities
none
Context window
128k tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.020
Output
$0.050
Cached Input
not available
Cache write
not available
Cost per task
not available

Capability

Intelligence
6.9
Coding
not available
Agentic
not available
Omniscience
-30.9
Correct
8.5%
Blended price
not available

Answer outcomes

Fully graded outcomes (Correct + Incorrect): 47.9%

Correct 8.5%Incorrect 39.4%Partial / not attempted 52.1%

Artificial Analysis benchmarks

GDPval-AA v2
0%
Humanity’s Last Exam
5%
GPQA Diamond
26%
CritPt
0%
AA-Omniscience
35%
AA-LCR
18%

Similar models

Support me! Patreon