readable AI benchmarks

NVIDIA

Llama 3.1 Nemotron Ultra 253B v1 (Reasoning)

reasoning · open weights · Apr 7, 2025

Quality Score14.1
Value Score10.5
Reliability27.6
Cache Discountnot available

Model specification

Reasoning
reasoning
Input modalities
text
Output modalities
text
Context window
128k tokens
Weights
open weights

Token prices USD per 1M tokens

Input
$0.600
Output
$1.80
Cached Input
not available
Cached Output
not available
Cost per task
not available

Capability

Intelligence
8.9
Coding
not available
Agentic
not available
Omniscience
-44.9
Correct
20.1%
Blended price
$0.720

Answer outcomes

Correct 20.1%Incorrect 64.9%Abstained 15.0%

Artificial Analysis benchmarks

SciCode
35%
Humanity’s Last Exam
7%
GPQA Diamond
73%
CritPt
0%
AA-Omniscience
28%
AA-LCR
8%

Similar models