This is an independent community listing. Check the source, license, data handling, and evaluation caveats before relying on a project. Inclusion is not an endorsement or security review.
Independent Jev-1.13.0 audit with reproducible code and per-call JSONL: tests abstention options, matched Korean and English items, question-shape interference, and option order; its strong abstention finding is on the KoBBQ dataset, not a universal calibration guarantee.
Early Python toolkit and CLINC150 study of conformal routing thresholds and prediction-powered audits over Jev 1.13 through OpenRouter, with request plans, 2,412 journalled answers and usage records, analysis code, and tests; its per-incoming-query risk bound depends on matching calibration and deployment traffic, and the study shows the scope gate missing its target after out-of-scope prevalence shifts while fallback accuracy remains unmeasured.
Reproducible OpenRouter measurements of Jev's response shapes, cost, and latency across eight use cases, plus a small head-to-head on 27 author-written support tickets; the author publishes raw data and corrections to earlier comparison errors.