This is an independent community listing. Check the source, license, data handling, and evaluation caveats before relying on a project. Inclusion is not an endorsement or security review.
Independent calibration measurement of Jev on two labelled datasets, Banking77 and Web of Science, with a Jev to frontier cascade priced per row from measured tokens, now also packaged as an installable tool (pip install janus-decide) that measures a threshold on your own data and ships none by default; the protocol was frozen before any result and the raw JSONL and figures are committed, and no routing parameter transferred between the two datasets, as the optimal threshold, the sign of the accuracy gap between the two models, and whether routing paid for itself all changed; the Web of Science labels come from publication metadata rather than per-document annotation, so part of the error measured there is label ambiguity.
Self-hosted GLiFormer 400M server for Choice, Score, and Noul through a Jev-compatible API, with public benchmark code and per-item results. Its documented JevBench comparison finds weaker accuracy than Jev on reasoning-heavy items; hosted cost figures are estimates and depend on deployment throughput.
Bilingual English and Chinese map of Jev use cases and failure modes that separates the author's small, raw-response API suites from cited third-party results and TypeSafe claims; its case studies help frame task-fit questions, but they do not establish population-level accuracy or a reusable confidence threshold.