Pytest-like framework for unit-testing LLM outputs with metrics for hallucination, relevancy, and bias.
✓ Live data verified
Self-host — open source, no signup or key; run it yourself and bring your own model keys.
DeepEval — Pytest-like framework for unit-testing LLM outputs with metrics for hallucination, relevancy, and bias.
Website: https://github.com/confident-ai/deepeval
Agent signup: self-host
Pricing: open-source
Full record: https://dreaming.press/api/tools/deepeval.jsonMachine record: /api/tools/deepeval.json
Evaluation toolkit for RAG pipelines — faithfulness, answer relevancy, and context metrics without ground truth.
Test-driven prompt and agent development — evals, red-teaming, and side-by-side model comparison from the CLI.
DeepEval is Pytest-like framework for unit-testing LLM outputs with metrics for hallucination, relevancy, and bias. It's in the Evals & testing category of the dreaming.press tool directory.
DeepEval is free and open source.
DeepEval does not publish an official MCP server as of our last check.
Yes — DeepEval is open source and self-hostable; you bring your own model keys.
Popular evals & testing alternatives to DeepEval include Ragas, promptfoo. Compare them in the dreaming.press directory.
We track the AI stack so you don't have to — pricing, MCP support, and which tools an agent can sign up for. Free.