JevBench
The open benchmark for Jev-class decision models. One bounded rubric in, one answer distribution out — measured on intelligence above chance, calibration, speed, and cost.
How the score works
JevBench Score (v1.4): harmonic mean of Intelligence, Calibration, Speed and Cost, 25% each. Fresh sealed items contribute 20% of Intelligence. Calibration blends toward the sealed-inclusive measure. Public-to-sealed gaps above 25 points reduce Intelligence. The low-Intelligence penalty applies below 50; Speed and Cost each have a separate Jev-class gate below 50. See the full method.
What changed in v1.4
Fresh sealed decisions counter public-set saturation. A public-to-sealed gap penalty rewards generalization. The harmonic score and Speed/Cost gates keep the board focused on fast, affordable Jev-class systems. An API flag shows when an operator endpoint received sealed item text, without answers. The private set will evolve to avoid leakage into training data.
Current leaderboard
Loading the live public results…
| Rank & system | Class | Score | Intelligence | Calibration | Speed | Cost | USD / 1k decisions | Licence |
|---|---|---|---|---|---|---|---|---|
| Loading current results… | ||||||||
Submit a system
Open an issue in the JevBench repository with a public endpoint or reproducible serving instructions. Mappings and evaluation conditions are documented before aggregation; operator endpoints may receive sealed item text without answers; those rows carry an API flag.