LM Evaluation Harness

Also known as
lm-evaluation-harness

What is LM Evaluation Harness?

A Python-based evaluation framework that enables testing of language models against 60+ standard academic benchmarks with support for various model formats, APIs, and custom evaluation metrics.

Independently observed

LM Evaluation Harness pricing

We don't have LM Evaluation Harness's full plan breakdown yet (its pricing page resisted automated reading). Here's what we could confirm. Always check live pricing for exact numbers.

Pricing as of verify at live pricing ↗Independently observed
Open sourceFree tier

Free and open source

What LM Evaluation Harness does

The capabilities that matter for ai evals testing, normalised so it lines up with every alternative. “-” means we haven't confirmed it, not that it's missing.

Capabilities
Architecture model
Open source CLI
LLM as a judge prompt grading framework
-
Specialized rag metrics faithfulness context relevance
-
Deterministic regex and json schema assertions
-
Synthetic test dataset generation from documents
-
Ci cd github actions pipeline blocking gates
-
Red teaming and adversarial vulnerability scanning
-
Multi model side by side ab regression testing
Human in the loop hitl annotation UI
-
Dashboard analytics for metric drift over time
-
SOC2 type ii
-
Mit or apache permissive oss license
-
Pricing model
Free open source
Independently observed

Platform & deployment

Independently observed
Platforms
  • CLI
Deployment
  • Self-hosted

Integrations (5)

Independently observed
  • Hugging Face
  • PyTorch
  • VLLM
  • GitHub
  • OpenAI-compliant APIs

Security & compliance

Known vulnerabilities: 0 (0 in the last 12 months) sourcea count reflects scale & disclosure, not quality

LM Evaluation Harness alternatives

Other ai evals testing we track, ranked by the same independent score.

All LM Evaluation Harness alternatives, ranked →

Compare LM Evaluation Harness

Side by side against other ai evals testing, attribute by attribute, with a source on every value.

Independent · unbought · dated

The Vioscale score: one lens on the evidence

Not user reviews and not a paid placement: a confidence-weighted blend of the independent signals below (adoption, activity, security posture, and more), which you can sort and re-weight yourself. Vendors can correct their listing but can never move their rank, and stars are weighted low as a vanity metric. It is one way to read the evidence for LM Evaluation Harness, not the verdict.

Balanced composite 54 / 100
low · 40%updating
Signal contributions to the composite score
SignalScoreWeightContributionEvidence
Pricing transparency800.086.7
Price level1000.055.2
Development activity540.095.1
Release cadence630.053.3
Dependent projects400.062.5
Capabilities290.082.4
Security score560.042.4
Integrations220.092.1
Stars780.032.0
Reliability00.070.0-
Security posture00.070.0-
Package downloads00.140.0-
Developer Q&A activity00.060.0-

Computed . Re-weight it by intent, or see the full method.

All data & sourcesshow ↓

Every value we hold, with its source, retrieval date, and confidence. This is the evidence behind the score: don't trust it, verify it.

Activity

AttributeValueEvidence
Commits last 30d50mediumsource · 2026-08-26 · 65%

Adoption

AttributeValueEvidence
Github stars13,802highsource · 2026-08-26 · 90%
Dependent repos252highsource · 2026-08-26 · 85%

Features

AttributeValueEvidence
CapabilitiesPricing model: free_open_source · Architecture model: open_source_cli · Multi model side by side ab regression testing: Yesmediumsource · 2026-08-21 · 60%

Integrations

AttributeValueEvidence
Count5mediumsource · 2026-08-21 · 60%

Language

AttributeValueEvidence
PrimaryPythonhighsource · 2026-08-26 · 90%

License

AttributeValueEvidence
SpdxMIThighsource · 2026-08-26 · 95%

Pricing

AttributeValueEvidence
Modelcommerciallowsource · 2026-08-26 · 48%
Free tierYesmediumsource · 2026-08-21 · 60%
Price levelfreemediumsource · 2026-08-21 · 60%
TransparentYesmediumsource · 2026-08-21 · 60%

Release

AttributeValueEvidence
Cadence days66mediumsource · 2026-08-26 · 70%
History18 itemsmediumsource · 2026-08-26 · 70%

Security

AttributeValueEvidence
Scorecard5.6highsource · 2026-08-26 · 90%
VulnerabilitiesCount: 0 · Source: https://advisories.ecosyste.ms/api/v1/advisories?ecosystem=pypi&package_name=lm-eval&per_page=100 · Last 12m: 0highsource · 2026-08-26 · 90%