# LM Evaluation Harness vs TruLens

| Attribute | LM Evaluation Harness | TruLens |
|---|---|---|
| **Vioscale score** | 53.7 (40% (low)) | 49.1 (32% (low)) |
| activity.commits_last_30d | 50 | 55 |
| adoption.dependent_repos | 252 | 1 |
| adoption.github_stars | 13,802 | 3,525 |
| deployment.options | `{"self_hosted":true}` | `{"self_hosted":true}` |
| description.long | A Python-based evaluation framework that enables testing of language models against 60+ standard academic benchmarks with support for various model formats, APIs, and custom evaluation metrics. | An open-source library for systematically evaluating and tracing LLM-based applications, providing feedback functions and metrics to assess quality, safety, and relevance. |
| features.capabilities | `{"pricing_model":"free_open_source","architecture_model":"open_source_cli","multi_model_side_by_side_ab_regression_testing":true}` | `{"pricing_model":"free_open_source","architecture_model":"open_source_cli","llm_as_a_judge_prompt_grading_framework":true,"dashboard_analytics_for_metric_drift_over_time":true,"multi_model_side_by_side_ab_regression_testing":true,"red_teaming_and_adversarial_vulnerability_scanning":true,"specialized_rag_metrics_faithfulness_context_relevance":true}` |
| integrations.count | 5 | 2 |
| integrations.list | `[{"name":"Hugging Face"},{"name":"PyTorch"},{"name":"VLLM"},{"name":"GitHub"},{"name":"OpenAI-compliant APIs"}]` | `[{"name":"OpenAI"},{"name":"GEPA"}]` |
| language.primary | Python | Python |
| license.spdx | MIT | MIT |
| platform.support | `{"cli":true}` | `{"cli":true,"web":true}` |
| pricing | `{"type":"open_source","summary":"Free and open source","freeTier":true,"sourceUrl":"https://github.com/pricing","retrievedAt":"2026-08-21T11:02:40.793Z"}` | `{"type":"open_source","freeTier":true,"sourceUrl":"https://www.trulens.org","retrievedAt":"2026-08-21T11:03:25.225Z"}` |
| pricing.free_tier | yes | yes |
| pricing.model | commercial | open_source |
| pricing.price_level | free | free |
| pricing.transparent | yes | yes |
| release.cadence_days | 66 | 14 |
| release.history | `[{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.12","date":"2026-05-11T13:04:19Z","type":"stable","version":"v0.4.12"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.11","date":"2026-02-13T20:21:47Z","type":"stable","version":"v0.4.11"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.10","date":"2026-01-27T19:56:53Z","type":"stable","version":"v0.4.10"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.9.2","date":"2025-11-26T23:27:06Z","type":"stable","version":"v0.4.9.2"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.9.1","date":"2025-08-04T11:36:05Z","type":"stable","version":"v0.4.9.1"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.9","date":"2025-06-19T14:18:27Z","type":"stable","version":"v0.4.9"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.8","date":"2025-03-05T07:49:46Z","type":"stable","version":"v0.4.8"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.7","date":"2024-12-17T10:37:09Z","type":"stable","version":"v0.4.7"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.6","date":"2024-11-25T13:38:29Z","type":"stable","version":"v0.4.6"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.5","date":"2024-10-08T21:06:05Z","type":"stable","version":"v0.4.5"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.4","date":"2024-09-05T15:13:13Z","type":"stable","version":"v0.4.4"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.3","date":"2024-07-01T14:00:36Z","type":"stable","version":"v0.4.3"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.2","date":"2024-03-18T13:07:28Z","type":"stable","version":"v0.4.2"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.1","date":"2024-01-31T15:29:14Z","type":"stable","version":"v0.4.1"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.4.0","date":"2023-12-04T15:08:53Z","type":"stable","version":"v0.4.0"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.3.0","date":"2022-12-08T08:34:37Z","type":"stable","version":"v0.3.0"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.2.0","date":"2022-03-07T02:12:23Z","type":"stable","version":"v0.2.0"},{"url":"https://github.com/EleutherAI/lm-evaluation-harness/releases/tag/v0.0.1","date":"2021-09-02T02:28:08Z","type":"stable","version":"v0.0.1"}]` | `[{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.13.1","date":"2026-08-20T18:47:11Z","type":"stable","version":"trulens-2.13.1"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.12.0","date":"2026-08-06T20:52:26Z","type":"stable","version":"trulens-2.12.0"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.11.0","date":"2026-08-04T17:22:33Z","type":"stable","version":"trulens-2.11.0"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.10.0","date":"2026-07-28T15:52:48Z","type":"stable","version":"trulens-2.10.0"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.9.0","date":"2026-07-23T15:33:25Z","type":"stable","version":"trulens-2.9.0"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.8.1","date":"2026-05-14T17:18:08Z","type":"stable","version":"trulens-2.8.1"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.8.0","date":"2026-04-30T19:16:57Z","type":"stable","version":"trulens-2.8.0"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.7.2","date":"2026-04-09T19:02:04Z","type":"stable","version":"trulens-2.7.2"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.7.1","date":"2026-03-10T19:15:51Z","type":"stable","version":"trulens-2.7.1"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.7.0","date":"2026-02-19T02:01:32Z","type":"stable","version":"trulens-2.7.0"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.6.0","date":"2026-02-04T20:34:59Z","type":"stable","version":"trulens-2.6.0"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.5.3","date":"2026-01-15T21:13:33Z","type":"stable","version":"trulens-2.5.3"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.5.2","date":"2025-12-11T23:00:50Z","type":"stable","version":"trulens-2.5.2"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.5.1","date":"2025-11-21T23:07:30Z","type":"stable","version":"trulens-2.5.1"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.5.0","date":"2025-11-14T23:56:50Z","type":"stable","version":"trulens-2.5.0"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.4.2","date":"2025-10-21T23:26:42Z","type":"stable","version":"trulens-2.4.2"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.4.1","date":"2025-10-08T07:24:27Z","type":"stable","version":"trulens-2.4.1"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.4.0","date":"2025-09-23T13:48:10Z","type":"stable","version":"trulens-2.4.0"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.3.1","date":"2025-09-04T20:26:49Z","type":"stable","version":"trulens-2.3.1"},{"url":"https://github.com/truera/trulens/releases/tag/trulens-2.3.0","date":"2025-08-28T20:33:52Z","type":"stable","version":"trulens-2.3.0"}]` |
| reliability.status_page | - | yes |
| security.scorecard | 5.6 | - |
| security.vulnerabilities | `{"count":0,"source":"https://advisories.ecosyste.ms/api/v1/advisories?ecosystem=pypi&package_name=lm-eval&per_page=100","last_12m":0,"max_severity":null}` | `{"count":0,"source":"https://advisories.ecosyste.ms/api/v1/advisories?ecosystem=pypi&package_name=trulens-eval&per_page=100","last_12m":0,"max_severity":null}` |

## Capabilities (AI Evals Testing)

| Capability | LM Evaluation Harness | TruLens |
|---|:--:|:--:|
| **Capabilities** |  |  |
| Architecture model | Open source CLI | Open source CLI |
| LLM as a judge prompt grading framework | - | ✓ |
| Specialized rag metrics faithfulness context relevance | - | ✓ |
| Deterministic regex and json schema assertions | - | - |
| Synthetic test dataset generation from documents | - | - |
| Ci cd github actions pipeline blocking gates | - | - |
| Red teaming and adversarial vulnerability scanning | - | ✓ |
| Multi model side by side ab regression testing | ✓ | ✓ |
| Human in the loop hitl annotation UI | - | - |
| Dashboard analytics for metric drift over time | - | ✓ |
| SOC2 type ii | - | - |
| Mit or apache permissive oss license | - | - |
| Pricing model | Free open source | Free open source |

*Source: Vioscale. Generated 2026-09-01T14:46:18.694Z. "-" = undocumented, not absent.*
