# LM Evaluation Harness vs Ragas

| Attribute | LM Evaluation Harness | Ragas |
|---|---|---|
| **Vioscale score** | 53.7 (40% (low)) | 58.4 (27% (low)) |
| activity.commits_last_30d | 50 | - |
| adoption.dependent_repos | 252 | 1 |
| adoption.github_stars | 13,802 | 15,484 |
| deployment.options | `{"self_hosted":true}` | `{"self_hosted":true}` |
| description.long | A Python-based evaluation framework that enables testing of language models against 60+ standard academic benchmarks with support for various model formats, APIs, and custom evaluation metrics. | A Python-based framework that evaluates RAG applications through automatic metrics covering faithfulness, relevance, and recall. Includes tools to synthetically generate test datasets customized for specific use cases, enabling developers to assess LLM application performance at both component and end-to-end levels. |
| features.capabilities | `{"pricing_model":"free_open_source","architecture_model":"open_source_cli","multi_model_side_by_side_ab_regression_testing":true}` | `{"open_source":true,"pricing_model":"free_open_source","architecture_model":"open_source_cli","built_in_rag_eval_metrics":true,"llm_as_a_judge_prompt_grading_framework":true,"synthetic_test_dataset_generation_from_documents":true}` |
| integrations.count | 5 | 3 |
| integrations.list | `[{"name":"Hugging Face"},{"name":"PyTorch"},{"name":"VLLM"},{"name":"GitHub"},{"name":"OpenAI-compliant APIs"}]` | `[{"name":"LlamaIndex"},{"name":"LangSmith"},{"name":"OpenAI"}]` |
| language.primary | Python | Python |
| license.spdx | MIT | Apache-2.0 |
| platform.support | `{"cli":true}` | `{"cli":true}` |
| pricing | `{"type":"open_source","summary":"Free and open source","freeTier":true,"sourceUrl":"https://github.com/pricing","retrievedAt":"2026-08-21T11:02:40.793Z"}` | `{"type":"open_source","freeTier":true,"sourceUrl":"https://ragas.io","retrievedAt":"2026-08-21T13:52:19.928Z"}` |
| pricing.free_tier | yes | yes |
| pricing.model | commercial | commercial |
| pricing.price_level | free | free |
| pricing.transparent | yes | yes |
| release.cadence_days | 66 | 7 |
| release.history | `[{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.12","date":"2026-05-11T13:04:19Z","type":"stable","version":"v0.4.12"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.11","date":"2026-02-13T20:21:47Z","type":"stable","version":"v0.4.11"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.10","date":"2026-01-27T19:56:53Z","type":"stable","version":"v0.4.10"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.9.2","date":"2025-11-26T23:27:06Z","type":"stable","version":"v0.4.9.2"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.9.1","date":"2025-08-04T11:36:05Z","type":"stable","version":"v0.4.9.1"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.9","date":"2025-06-19T14:18:27Z","type":"stable","version":"v0.4.9"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.8","date":"2025-03-05T07:49:46Z","type":"stable","version":"v0.4.8"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.7","date":"2024-12-17T10:37:09Z","type":"stable","version":"v0.4.7"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.6","date":"2024-11-25T13:38:29Z","type":"stable","version":"v0.4.6"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.5","date":"2024-10-08T21:06:05Z","type":"stable","version":"v0.4.5"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.4","date":"2024-09-05T15:13:13Z","type":"stable","version":"v0.4.4"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.3","date":"2024-07-01T14:00:36Z","type":"stable","version":"v0.4.3"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.2","date":"2024-03-18T13:07:28Z","type":"stable","version":"v0.4.2"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.1","date":"2024-01-31T15:29:14Z","type":"stable","version":"v0.4.1"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.0","date":"2023-12-04T15:08:53Z","type":"stable","version":"v0.4.0"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.3.0","date":"2022-12-08T08:34:37Z","type":"stable","version":"v0.3.0"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.2.0","date":"2022-03-07T02:12:23Z","type":"stable","version":"v0.2.0"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.0.1","date":"2021-09-02T02:28:08Z","type":"stable","version":"v0.0.1"}]` | `[{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.4.3","date":"2026-01-13T17:47:29Z","type":"stable","version":"v0.4.3"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.4.2","date":"2025-12-23T17:13:41Z","type":"stable","version":"v0.4.2"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.4.1","date":"2025-12-10T16:28:51Z","type":"stable","version":"v0.4.1"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.4.0","date":"2025-12-03T16:22:29Z","type":"stable","version":"v0.4.0"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.3.9","date":"2025-11-11T17:24:46Z","type":"stable","version":"v0.3.9"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.3.8","date":"2025-10-28T19:09:17Z","type":"stable","version":"v0.3.8"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.3.7","date":"2025-10-14T16:21:37Z","type":"stable","version":"v0.3.7"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.3.6","date":"2025-10-03T03:56:31Z","type":"stable","version":"v0.3.6"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.3.5","date":"2025-09-17T19:13:15Z","type":"stable","version":"v0.3.5"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.3.5rc2","date":"2025-09-17T17:40:25Z","type":"prerelease","version":"v0.3.5rc2"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.3.5rc1","date":"2025-09-17T17:31:23Z","type":"prerelease","version":"v0.3.5rc1"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.3.4","date":"2025-09-10T23:57:28Z","type":"stable","version":"v0.3.4"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.3.3","date":"2025-09-04T17:59:51Z","type":"stable","version":"v0.3.3"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.3.3rc1","date":"2025-09-04T17:46:09Z","type":"prerelease","version":"v0.3.3rc1"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.3.2","date":"2025-08-19T12:03:50Z","type":"stable","version":"v0.3.2"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.3.2rc3","date":"2025-08-19T12:01:12Z","type":"prerelease","version":"v0.3.2rc3"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.3.2-rc2","date":"2025-08-19T11:49:23Z","type":"prerelease","version":"v0.3.2-rc2"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.3.2-rc1","date":"2025-08-19T10:28:33Z","type":"prerelease","version":"v0.3.2-rc1"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.3.1","date":"2025-08-11T11:15:03Z","type":"stable","version":"v0.3.1"},{"url":"https://github.com/vibrantlabsai/ragas/releases/tag/v0.3.0","date":"2025-07-17T05:32:24Z","type":"stable","version":"v0.3.0"}]` |
| security.scorecard | 5.6 | - |
| security.vulnerabilities | `{"count":0,"source":"https://advisories.ecosyste.ms/api/v1/advisories?ecosystem=pypi&package_name=lm-eval&per_page=100","last_12m":0,"max_severity":null}` | `{"count":2,"source":"https://advisories.ecosyste.ms/api/v1/advisories?ecosystem=pypi&package_name=ragas&per_page=100","last_12m":2,"max_severity":"HIGH"}` |

## Capabilities (AI Evals Testing)

| Capability | LM Evaluation Harness | Ragas |
|---|:--:|:--:|
| **Capabilities** |  |  |
| Architecture model | Open source CLI | Open source CLI |
| LLM as a judge prompt grading framework | - | ✓ |
| Specialized rag metrics faithfulness context relevance | - | - |
| Deterministic regex and json schema assertions | - | - |
| Synthetic test dataset generation from documents | - | ✓ |
| Ci cd github actions pipeline blocking gates | - | - |
| Red teaming and adversarial vulnerability scanning | - | - |
| Multi model side by side ab regression testing | ✓ | - |
| Human in the loop hitl annotation UI | - | - |
| Dashboard analytics for metric drift over time | - | - |
| SOC2 type ii | - | - |
| Mit or apache permissive oss license | - | - |
| Pricing model | Free open source | Free open source |

*Source: Vioscale. Generated 2026-09-01T16:46:53.219Z. "-" = undocumented, not absent.*
