Skip to content
#

agent-evaluation

Here are 1,224 public repositories matching this topic...

Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.

  • Updated Oct 5, 2026
  • Python

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.

  • Updated Oct 5, 2026
  • Python

OpenART is an open-source framework designed to evaluate the safety and robustness of autonomous AI agents in dynamic, long-horizon, and stateful environments. It stress-tests agent runtimes against multi-step state poisoning, privilege escalation, and tool-use vulnerabilities across 10,000+ benchmark scenarios.

  • Updated Oct 3, 2026
  • Python
AgentMeasure

Independent verification of AI support bills: applicable terms, invoice and business records, with reviewable findings. Open-source conformance for agent telemetry — PASS / FAIL / UNPROVABLE in CI.

  • Updated Oct 5, 2026
  • Python
coder_eval

Playwright for coding agents. Test that your skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, A/B experiments, CI gates.

  • Updated Oct 5, 2026
  • Python

Add this topic to your repo

To associate your repository with the agent-evaluation topic, visit your repo's landing page and select "manage topics."

Learn more