Type

Tag “llm evaluation” — 7

Langfuse

langfuse.com

LLM engineering platform for observability and metrics

Agent Infrastructure & MLOps Web + desktop

LangWatch

langwatch.ai

Test and improve your AI agents before users find issues.

Agent Infrastructure & MLOps Web only

Agenta

agenta.ai

Agenta helps teams build, test, and manage reliable AI applications.

Agent Infrastructure & MLOps Command line

ModelBench

modelbench.ai

Test and compare AI models easily, with no code. Launch faster.

Agent Infrastructure & MLOps Web only

Evidently AI

evidentlyai.com

Open-source и облачная платформа для оценки, тестирования и мониторинга AI- и ML-моделей с обширными метриками и инструментами для совместной работы.

Tool Coding & development Self-hosted

Confident AI

confident-ai.com

Automate AI testing and ensure your AI apps always work well.

Agent Infrastructure & MLOps Command line

Phoenix

phoenix.arize.com

Phoenix Arize is an open-source AI observability platform for LLMs.

Agent Data & analytics Command line

Submit a site to the catalog

Just send the link — we will work out the rest.

We will review what you send and add it to the catalog if it fits.

Not sure how to implement it? We can help

Tell us about your task — we will pick the tools and suggest where to start.

0 / 5000
Verification code

Fields marked with an asterisk are required. Your data is used only to reply.