Rust example that asks a configured OpenRouter decision model to score which tool outputs are worth keeping during context compaction; it compares the ordering with an age-based policy on a synthetic session, not a live-agent outcome.
This is an independent community listing. Check the source, license, data handling, and evaluation caveats before relying on a project. Inclusion is not an endorsement or security review.
MIT TypeScript tool-call firewall for Claude Code (PreToolUse hook) and MCP servers (stdio proxy): static rules first, then seven Jev Noul risk questions per action mapped to allow, ask, or deny by policy thresholds, with an action ledger, key-aware redaction, and a local JSONL audit log. Sends redacted tool input, cwd, and recent prompts to TypeSafe, OpenRouter, or Vercel AI Gateway (owner's key); in one author's 1,369-decision window on 0.9.1, blind second labeling of 150 sampled allows found no permissive misses, while 21% of decisions required an ask, mostly from one axis. Version 0.11.0 addresses that axis but has not been re-measured. Heuristic redaction and model judgments are defense in depth, not a sandbox.
Agent framework whose auto model router defaults to Jev through Vercel AI Gateway and whose evaluate helper asks typed Choice, Score, and Boolean questions inside tools; the underlying AI SDK evaluation model specification is experimental, and the caller owns routing and effects.