Community project / Evaluations and independent research
Love-Language Arena
Experimental local reproduction of the Choice, Score, and Noul pattern on Ollama logprobs, with order-reversed and negated re-asks to expose position bias, used to test three open models as fully crossed judges of Chinese and English rewrites, with self-preference correction, per-judge Platt calibration, dev/holdout prompt selection, and raw per-call JSONL. It does not call Jev, and its validation labels are Claude-generated rather than human, so the judge pass rates and model ranking show that the pipeline runs, not that the judges are valid.
Download share card
Share this listing
Maintaining this project? Share its link, badge, or image. Listed means included, not endorsed.
This is an independent community listing. Check the source, license, data handling, and evaluation caveats before relying on a project. Inclusion is not an endorsement or security review.
Found an outdated or inaccurate detail? Report a correction →