Promptfoo
Automated testing platform that uses red teaming to identify and fix security vulnerabilities in AI applications before deployment
- Also known as
- promptfoo
Available worldwide · Popular in: US
What is Promptfoo?
A comprehensive LLM security testing solution that automatically scans for vulnerabilities such as prompt injections and jailbreaks through dynamic red teaming. Offers both open-source software for local testing and enterprise SaaS with team collaboration, continuous monitoring, and compliance dashboards.
Promptfoo pricing
Plans, per-tier features and add-ons, dated and linked to live pricing. Pricing changes often; always verify at source before you rely on it.
Free open-source Community tier; Enterprise and On-Premise plans with custom pricing
Community
FreeOpen-source version for individual developers and small teams
- red_teaming_probes
- 10k/mo
- All LLM evaluation features
- All model providers and integrations
- Red teaming (10k probes/month)
- Custom integration with your own app
- Run locally or self-host on your own infrastructure
- Vulnerability scanning
- Community support
Enterprise
Contact salesManaged cloud platform with team collaboration, continuous monitoring, and compliance features
- All Community features
- Custom red teaming limits
- Team sharing & collaboration
- Continuous monitoring
- Centralized security/compliance dashboard
- Customizable attack profiles and target settings
- SSO and granular permission profiles
- Promptfoo API access
- Managed cloud deployment
- Professional services support
- Priority support & SLA guarantees
On-Premise
Contact salesCustom deployment for organizations requiring full control over infrastructure
- Complete data isolation
- On-premise deployment
What Promptfoo does
The capabilities that matter for prompt management, normalised so it lines up with every alternative. “-” means we haven't confirmed it, not that it's missing.
- Versioning model
- UI hosted registry
- Diff and rollback
- -
- Deploy without redeploy
- -
- Non engineer editing
- ✗
- Multi model playground
- -
- Ab testing
- -
- LLM judge or human eval
- ✓
- Multi step flows
- -
- Approval workflow
- -
- Tracing included
- -
- Self hostable
- ✓
- Deployment model
- -
- Pricing model
- -
Platform & deployment
Independently observed- CLI
- Web
- Cloud / SaaS
- Hybrid
- On-premise
- Self-hosted
Integrations (11)
Independently observed- OpenAI
- Anthropic
- Google (Gemini)
- DeepSeek
- MCP (Model Context Protocol)
- CI/CD pipelines
- Webhooks
- Agent frameworks
- CI/CD platforms
- GitHub
- Python providers
Promptfoo alternatives
Other prompt management we track, ranked by the same independent score.
- PromptHubA collaborative platform for managing, testing, and deploying AI prompts across teamslow · 26%
- Orq.aiA unified AI gateway routing requests across 500+ models from multiple providers with built-in observability and governancemedium · 55%
- PromptLayerA collaboration layer enabling AI engineering teams to manage and improve prompts without modifying codemedium · 54%
- LatitudeAutomatically detect and repair AI agent failures in productionmedium · 54%
- VellumA personalized AI assistant with persistent memory that integrates across your digital life and devices.low · 37%
- AgentaPlatform for building, testing, and deploying AI agents with team collaboration and iterative refinementlow · 43%
The Vioscale score: one lens on the evidence
Not user reviews and not a paid placement: a confidence-weighted blend of the independent signals below (adoption, activity, security posture, and more), which you can sort and re-weight yourself. Vendors can correct their listing but can never move their rank, and stars are weighted low as a vanity metric. It is one way to read the evidence for Promptfoo, not the verdict.
| Signal | Score | Weight | Contribution | Evidence |
|---|---|---|---|---|
| Security posture | 65 | 0.07 | 4.7 | ✓ |
| Price level | 80 | 0.05 | 4.2 | ✓ |
| Capabilities | 87 | 0.05 | 4.2 | ✓ |
| Pricing transparency | 25 | 0.08 | 2.1 | ✓ |
| Integrations | 24 | 0.04 | 1.0 | ✓ |
| Reliability | 0 | 0.07 | 0.0 | - |
Computed . Re-weight it by intent, or see the full method.
All data & sourcesshow ↓
Every value we hold, with its source, retrieval date, and confidence. This is the evidence behind the score: don't trust it, verify it.
Features
| Attribute | Value | Evidence |
|---|---|---|
| Capabilities | Soc2 type ii: Yes · Self hostable: Yes · Versioning model: UI-hosted registry · Non engineer editing: No · Llm judge or human eval: Yes · Llm as a judge prompt grading framework: Yes | mediumsource · 2026-08-21 · 60% |
Integrations
| Attribute | Value | Evidence |
|---|---|---|
| Count | 6 | mediumsource · 2026-08-21 · 60% |
Market
| Attribute | Value | Evidence |
|---|---|---|
| Availability | HqCountry: US · PrimaryMarkets: … · AvailabilityScope: global · AvailableCountries: … · NotAvailableCountries: … | highsource · 2026-08-21 · 75% |