Community project / Evaluations and independent research
Jev Judge vs Dimension Scores
Independent measurement on three classification tasks: one direct Jev question per row against 12–14 Jev-scored dimensions with locally fitted weights, 5,477 test rows and 34.1M input tokens for $1.43; decomposition reached 0.9076 against 0.8373 on Japanese NLI but flagged about 25× more hard benign rows as attacks, and four repair attempts failed, on dimensions the author wrote himself.
Download share card
Share this listing
Maintaining this project? Share its link, badge, or image. Listed means included, not endorsed.
This is an independent community listing. Check the source, license, data handling, and evaluation caveats before relying on a project. Inclusion is not an endorsement or security review.
Found an outdated or inaccurate detail? Report a correction →