readable AI benchmarks (simplified)

readable AI benchmarks

Anthropic

Claude 3.5 Sonnet (June '24)

non-reasoning · closed weights · Jun 21, 2024

Claude 3.5 Sonnet (June '24) does not yet have enough benchmark data for a Quality Score. Per 1M tokens, pricing is $3.00 input, $15.00 output, $0.300 cached input, and $3.75 cache write. It has a 200k token context window.

Quality Scorenot availableQuality unavailable: missing AA-Omniscience outcomes
Value Scorenot available
Factual reliabilitynot available
Cache Discount90%

Model specification

Reasoning
non-reasoning
Input modalities
none
Output modalities
none
Context window
200k tokens
Weights
closed weights

Token prices USD per 1M tokens

Input
$3.00
Output
$15.00
Cached Input
$0.300
Cache write
$3.75
Cost per task
not available

Capability

Intelligence
7.2
Coding
not available
Agentic
not available
Omniscience
not available
Correct
not available
Blended price
$2.31

Answer outcomes

Fully graded outcomes (Correct + Incorrect): not available

Omniscience outcome data not available for this model.

Artificial Analysis benchmarks

Humanity’s Last Exam
3%
GPQA Diamond
56%

Similar models

Support me! Patreon