← Explore
TOPIC

#evaluation

Open source repositories tagged with #evaluation, ranked by health score.

mlflow
mlflow/mlflow
Python
89
health

The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.

★ 28.3k
comet-ml
comet-ml/opik
Python
89
health

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.

★ 22.4k
promptfoo
promptfoo/promptfoo
TypeScript
88
health

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

★ 25.7k