Community project / Evaluations and independent research
jev-ood-calibration
Independent Jev calibration study with raw Gateway responses from three public benchmarks and 900 rule-generated support tickets; probabilities were near calibrated on the public sets but overconfident on an unseen priority rule, while the Boolean question showed a different error direction. The synthetic task is one family and the Gateway did not expose a fixed model version.
This is an independent community listing. Check the source, license, data handling, and evaluation caveats before relying on a project. Inclusion is not an endorsement or security review.
Independent measurement of whether ORDER BY over a Jev probability is defensible, with a pre-registered gate on pairwise inversion, Score ordinality against a graded target, calibration, and negation and paraphrase invariants; jev-1.13.0 passes on 20 Newsgroups topic membership and fails four of six conditions on Amazon ESCI human-graded product relevance, the DuckDB integrations' request shapes are shown to change the numbers (a 40-row batched state fails the ranking gate that one row per request passes), and two-decimal output leaves 53 of 360 rows tied at the top so LIMIT k cuts inside a tie; one seed and 30 ESCI queries, aggregates committed, corpus and cached responses regenerated locally for about a cent.
Blind Jev-1.13.0 security evaluation with Go runner, raw per-sample results, and a TUI: 662 public prompt-injection messages and 200 matched vulnerable-code pairs; its reported classification scores use the study's stated context and a fixed 0.5 threshold, so deployment policy and corpus labels matter.
Independent benchmark of hosted Jev and six LLMs on PubMedQA, Banking77, and HelpSteer2, with human labels, proper scores, calibration, cost, and latency; suite files and per-decision logs and the methodology are public, while the run harness is not published, LLM probabilities come from a prompt adapter, and the Gateway does not expose Jev's exact version.