← Explore
TOPIC

#llm-evaluation

Open source repositories tagged with #llm-evaluation, ranked by health score.

ifixai-ai
ifixai-ai/iFixAi
Python
89
health

Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.

★ 22.4k
mlflow
mlflow/mlflow
Python
89
health

The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while controlling costs and managing access to models and data.

★ 28.3k
Tencent
Tencent/AI-Infra-Guard
Python
89
health

A full-stack AI Red Teaming platform securing AI ecosystems via Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.

★ 6.8k
scaleapi
scaleapi/agentenv-framework
Python
87
health

Creating realistic RL environments requires collaboration between researchers, engineers, and domain experts across many dimensions: artifacts, environments tools, dynamism of the environment, reproducibility, and more. There is no open source framework for building these environments effectively. Until now.

★ 189
memex-lab
memex-lab/dart_agent_core
Dart
83
health

Dart framework for stateful AI agents: tool use, skills, sub-agent delegation, planning, streaming, evals, and multi-provider LLM support.

★ 46
cookiespiggy
cookiespiggy/agentic-rl
Python
80
health

Agentic RL 中文零基础教程(33 章):从概念到 GRPO 实战,含 TRL 最小可跑示例;26–33 章附一套可运行的三方判别模型实证工程(encoder vs LLM-LoRA vs 规则基线)。第 25 章讲清 Jev / TypeSafe System One 与 RL 的能力边界 | Chinese Agentic RL tutorial (33 chapters) + a reproducible discriminative-model benchmark

★ 119