# OpenAI Evals

- **Canonical URI:** https://www.vioscale.ai/software/openai-evals
- **Category:** AI Evals Testing
- **Homepage:** https://github.com/openai/evals
- **Also known as:** openai-evals
- **Profile claimed by vendor:** no
- **Last updated:** 2026-08-26T19:34:57.687Z

## Vioscale score

**44.8 / 100**, confidence 30% (low). Computed 2026-09-01.

Composite of weighted, independently-sourced signals (no user reviews, no vendor payment).

| Signal | Score | Weight | Contribution | Evidence present |
|---|--:|--:|--:|:--:|
| price_level | 100 | 0.052 | 5.2 | ✓ |
| reliability | 0 | 0.073 | 0 | - |
| capabilities | 74.7 | 0.084 | 6.3 | ✓ |
| repo_stars | 80.8 | 0.026 | 2.1 | ✓ |
| integrations | 17.3 | 0.09355555555555557 | 1.6 | ✓ |
| dependent_projects | 5 | 0.063 | 0.3 | ✓ |
| dev_activity | 0 | 0.094 | 0 | ✓ |
| release_cadence | 0 | 0.052 | 0 | - |
| security_posture | 5 | 0.073 | 0 | - |
| package_downloads | 0 | 0.136 | 0 | - |
| security_score | 0 | 0.042 | 0 | - |
| pricing_transparency | 80 | 0.084 | 6.7 | ✓ |
| community_qa_activity | 0 | 0.063 | 0 | - |

## Pricing

_As of 2026-08-21, [verify at source](https://github.com/pricing). Independently observed._

Free · Free tier

> Free and open-source

## About

A framework that lets developers create and run evaluations to measure LLM performance, providing both pre-built benchmarks and tools to write custom tests tailored to specific use cases without requiring proprietary evaluation infrastructure.

_Independently observed._

## Platform & deployment

- **Platforms:** CLI, Web
- **Deployment:** Cloud / SaaS, Self-hosted

## Integrations (3)

_Independently observed._

- OpenAI API
- Snowflake
- GitHub

## Capabilities

_The capabilities that matter for AI Evals Testing. "-" = undocumented, not absent._

| Capability | Supported |
|---|:--:|
| **Capabilities** | |
| Architecture model | Open source CLI |
| LLM as a judge prompt grading framework | ✓ |
| Specialized rag metrics faithfulness context relevance | - |
| Deterministic regex and json schema assertions | ✓ |
| Synthetic test dataset generation from documents | ✓ |
| Ci cd github actions pipeline blocking gates | ✓ |
| Red teaming and adversarial vulnerability scanning | - |
| Multi model side by side ab regression testing | - |
| Human in the loop hitl annotation UI | - |
| Dashboard analytics for metric drift over time | - |
| SOC2 type ii | - |
| Mit or apache permissive oss license | ✓ |
| Pricing model | Free open source |

## Security & compliance


Known vulnerabilities: 0 (0 in the last 12 months) ([source](https://advisories.ecosyste.ms/api/v1/advisories?ecosystem=go&package_name=github.com%2Fopenai%2Fevals&per_page=100)). A count reflects scale and disclosure, not quality.

## Facts

Every value below carries its source and our confidence. Facts are re-crawled on a freshness schedule.

### pricing

| Attribute | Value | Source | Retrieved | Confidence |
|---|---|---|---|---|
| pricing.model | commercial | [link](https://github.com/openai/evals) | 2026-08-26 | 60% (medium) |
| pricing.price_level | free | [link](https://github.com/pricing) | 2026-08-21 | 60% (medium) |
| pricing.transparent | yes | [link](https://github.com/pricing) | 2026-08-21 | 60% (medium) |
| pricing.free_tier | yes | [link](https://github.com/pricing) | 2026-08-21 | 60% (medium) |

### security

| Attribute | Value | Source | Retrieved | Confidence |
|---|---|---|---|---|
| security.disclosure_policy | yes | [link](https://github.com/pricing) | 2026-08-21 | 60% (medium) |
| security.vulnerabilities | `{"count":0,"source":"https://advisories.ecosyste.ms/api/v1/advisories?ecosystem=go&package_name=github.com%2Fopenai%2Fevals&per_page=100","last_12m":0,"max_severity":null}` | [link](https://advisories.ecosyste.ms/api/v1/advisories?ecosystem=go&package_name=github.com%2Fopenai%2Fevals&per_page=100) | 2026-08-26 | 90% (high) |

### integrations

| Attribute | Value | Source | Retrieved | Confidence |
|---|---|---|---|---|
| integrations.count | 3 | [link](https://github.com/pricing) | 2026-08-21 | 60% (medium) |

### activity

| Attribute | Value | Source | Retrieved | Confidence |
|---|---|---|---|---|
| activity.commits_last_30d | 0 | [link](https://github.com/openai/evals/pulse) | 2026-08-26 | 65% (medium) |

### adoption

| Attribute | Value | Source | Retrieved | Confidence |
|---|---|---|---|---|
| adoption.github_stars | 19,257 | [link](https://github.com/openai/evals) | 2026-08-26 | 90% (high) |
| adoption.dependent_repos | 1 | [link](https://packages.ecosyste.ms/api/v1/packages/lookup?repository_url=https%3A%2F%2Fgithub.com%2Fopenai%2Fevals) | 2026-08-26 | 85% (high) |

### language

| Attribute | Value | Source | Retrieved | Confidence |
|---|---|---|---|---|
| language.primary | Python | [link](https://github.com/openai/evals) | 2026-08-26 | 90% (high) |

---
*Source: Vioscale (https://www.vioscale.ai/software/openai-evals). Independent, evidence-based software intelligence. Cite the canonical URI.*
