Evaluation

Measure whether output is actually good — faithfulness, relevance, correctness — with metrics and test suites you can run in CI.

Pricing
License
Language
Last updated: Show all
Sort: A–Z
DeepEval

DeepEval

Confident AI

Open-source "Pytest for LLMs" — unit-test your model and RAG output with metrics.

v4.1.32026-07-12PythonApache-2.0
17.1k Open source Free tier Paid
Ragas

Ragas

Exploding Gradients

Open-source evaluation framework for RAG pipelines and LLM applications

No updates in 6+ monthsv0.4.32026-01-13PythonApache-2.0
15.0k Completely free Open source