Measuring agent and LLM output quality, regressions, and safety. Ranked by community traction, with live GitHub stars and what each is best at.
Short answer: the best evals & testing for AI agents by community traction is promptfoo (★ 24k), followed by DeepEval and Ragas.
✓ Live data verified
Test-driven prompt and agent development — evals, red-teaming, and side-by-side model comparison from the CLI. Best for prompt evals.
Pytest-like framework for unit-testing LLM outputs with metrics for hallucination, relevancy, and bias. Best for LLM unit tests.
Evaluation toolkit for RAG pipelines — faithfulness, answer relevancy, and context metrics without ground truth. Best for RAG evaluation.
By community traction, promptfoo (★ 24k) leads the evals & testing in our directory. Test-driven prompt and agent development — evals, red-teaming, and side-by-side model comparison from the CLI.
promptfoo is the most-starred open-source option; DeepEval and Ragas are strong runners-up.
promptfoo, at ★ 24k (live count).
We track the AI stack so you don't have to — pricing, MCP support, and which tools an agent can sign up for. Free.