Community project / Evaluations and independent research
JevBench
Independent cross-model benchmark for typed decisions with 534 frozen cases per complete entrant, source adapters, scoring code, public per-task outcomes, and aggregate result artifacts. Its four-axis score includes accuracy, calibration, latency, and cost; some self-hosted latency is adjusted by assumption, hosting costs can be estimates, and withheld cases are exposed to the services being measured.
Download share card
Share this listing
Maintaining this project? Share its link, badge, or image. Listed means included, not endorsed.
This is an independent community listing. Check the source, license, data handling, and evaluation caveats before relying on a project. Inclusion is not an endorsement or security review.
Found an outdated or inaccurate detail? Report a correction →