Benchmark Heaven
Open benchmark

JevBench

The open benchmark for Jev-class decision models. One bounded rubric in, one answer distribution out — measured on intelligence above chance, calibration, speed, and cost.

Explore the full interactive board Code & method Submit a system
ranked systems
842decisions per complete run
4equally weighted axes
current revision

How the score works

IntelligenceChance-corrected accuracy on the frozen public tiers, blended with 308 fresh sealed decisions.
CalibrationDistribution quality on public items, blended toward calibration that includes sealed items.
SpeedSerial p50 and p95 latency on a logarithmic scale, self-hosted endpoints adjusted for production load.
CostUSD per 1,000 complete decisions — a whole question, not per 1,000 tokens.

JevBench Score (v1.4): harmonic mean of Intelligence, Calibration, Speed and Cost, 25% each. Fresh sealed items contribute 20% of Intelligence. Calibration blends toward the sealed-inclusive measure. Public-to-sealed gaps above 25 points reduce Intelligence. The low-Intelligence penalty applies below 50; Speed and Cost each have a separate Jev-class gate below 50. See the full method.

What changed in v1.4

Fresh sealed decisions counter public-set saturation. A public-to-sealed gap penalty rewards generalization. The harmonic score and Speed/Cost gates keep the board focused on fast, affordable Jev-class systems. An API flag shows when an operator endpoint received sealed item text, without answers. The private set will evolve to avoid leakage into training data.

Current leaderboard

Loading the live public results…

Rank & system Class Score Intelligence Calibration Speed Cost USD / 1k decisions Licence
Loading current results…

Submit a system

Open an issue in the JevBench repository with a public endpoint or reproducible serving instructions. Mappings and evaluation conditions are documented before aggregation; operator endpoints may receive sealed item text without answers; those rows carry an API flag.