Open-source LLM and agent evaluation framework with 50+ metrics, LLM-as-Judge augmentation, and guardrail scanners (jailbreak, PII, prompt-injection). Useful for scoring RAG outputs, agent trajectories, and function-calling behavior in data-science workflows.
| Type | Repository |
| Section | GitHub projects |
| Pricing | open source |
| Platform | Self-hosted |
| Systems | исходный код |
| Site language | en |
| GitHub | future-agi/ai-evaluation |