Community project / Evaluations and independent research

JevBench

Independent cross-model benchmark for typed decisions with 534 frozen cases per complete entrant, source adapters, scoring code, public per-task outcomes, and aggregate result artifacts. Its four-axis score includes accuracy, calibration, latency, and cost; some self-hosted latency is adjusted by assumption, hosting costs can be estimates, and withheld cases are exposed to the services being measured.

Awesome Jev share card for JevBench Download share card

Share this listing

Maintaining this project? Share its link, badge, or image. Listed means included, not endorsed.

This is an independent community listing. Check the source, license, data handling, and evaluation caveats before relying on a project. Inclusion is not an endorsement or security review.

Found an outdated or inaccurate detail? Report a correction →

Explore more in Evaluations and independent research