Community project / Client libraries and integrations
grounded-ai
MIT Python evaluation library whose JevEvaluator asks hosted Jev or a local Strands Decider Noul, Choice, and Score questions, and whose CascadeEvaluator keeps confident answers and sends the rest to an LLM judge (OpenAI, Anthropic, Bedrock, or Hugging Face). Hosted mode sends inputs to TypeSafe; escalated questions go to the judge's provider even when the first judge is local. Its tool-call benchmark has 200 items labelled by another model's choices.
Download share card
Share this listing
Maintaining this project? Share its link, badge, or image. Listed means included, not endorsed.
This is an independent community listing. Check the source, license, data handling, and evaluation caveats before relying on a project. Inclusion is not an endorsement or security review.
Found an outdated or inaccurate detail? Report a correction →