DeepEval
Confident AIPytest-style evaluation framework for LLM apps with 40+ built-in metrics including G-Eval, hallucination, faithfulness and RAG-specific scores. Picked for CI/CD integration; drops into existing test suites and gates merges on metric thresholds. 15.6k stars, 263 contributors, 561 commits in last 90 days, the most active open source eval framework on GitHub.
DeepEval v4.0.5 ('Opus 4.8 Day-0 Support', May 28 2026). OSS eval framework (Apache-2.0) with commercial Confident AI cloud platform. Confirmed live June 2026.
Openness
4 high confidence- license
- Apache-2.0(OSI, OSS framework)
- source
- public(GitHub)
- commercial-tier
- Confident AI managed eval/observability cloud on top
- core-gated
- gated
Apache-2.0 OSS core + proprietary Confident AI cloud -> open_core (recipe names Confident AI as the open_core exemplar); the library itself is OSI.
- https://github.com/confident-ai/deepeval recorded 2026-06-04
Apache-2.0 license; Confident AI commercial cloud platform
Adoption
4 high confidence~4.63M monthly PyPI downloads (deepeval); 15.9k stars, used-by 1.3k projects. Headline = PyPI primary source; download volume places it firmly in the 1-10M band.
- https://pypistats.org/packages/deepeval recorded 2026-06-04
~4,628,371 monthly PyPI downloads
- https://github.com/confident-ai/deepeval recorded 2026-06-04
15.9k stars; used by 1.3k projects
Capability
4 high confidenceStrong, broad app/RAG/agentic eval coverage with LLM-as-judge -- top-tier for product-eval, just below the academic-benchmark standards (lm-eval) on benchmark-standardization breadth.
- https://github.com/confident-ai/deepeval recorded 2026-06-04
G-Eval, RAG/agentic/conversational metrics, pytest-style eval
Unchanged since 2026-07-30 (last edited, not re-checked)