Galileo
AI evaluation and observability platform for testing, monitoring, and improving AI systems
- Also known as
- galileo
What is Galileo?
A platform for building and evaluating AI systems that synthesizes test datasets from multiple sources, compresses expensive LLM evaluators into efficient models, and provides production monitoring with real-time observability.
What Galileo does
The capabilities that matter for ai evals testing, normalised so it lines up with every alternative. “-” means we haven't confirmed it, not that it's missing.
- Architecture model
- Managed enterprise saas
- LLM as a judge prompt grading framework
- ✓
- Specialized rag metrics faithfulness context relevance
- -
- Deterministic regex and json schema assertions
- -
- Synthetic test dataset generation from documents
- ✓
- Ci cd github actions pipeline blocking gates
- ✓
- Red teaming and adversarial vulnerability scanning
- -
- Multi model side by side ab regression testing
- -
- Human in the loop hitl annotation UI
- ✓
- Dashboard analytics for metric drift over time
- ✓
- SOC2 type ii
- -
- Mit or apache permissive oss license
- -
- Pricing model
- -
Integrations (1)
Independently observed- NVIDIA NeMo
Galileo alternatives
Other ai evals testing we track, ranked by the same independent score.
The Vioscale score: one lens on the evidence
Not user reviews and not a paid placement: a confidence-weighted blend of the independent signals below (adoption, activity, security posture, and more), which you can sort and re-weight yourself. Vendors can correct their listing but can never move their rank, and stars are weighted low as a vanity metric. It is one way to read the evidence for Galileo, not the verdict.
| Signal | Score | Weight | Contribution | Evidence |
|---|---|---|---|---|
| Capabilities | 75 | 0.05 | 3.6 | ✓ |
| Integrations | 9 | 0.04 | 0.4 | ✓ |
| Price level | 0 | 0.05 | 0.0 | - |
| Reliability | 0 | 0.07 | 0.0 | - |
| Security posture | 0 | 0.07 | 0.0 | - |
| Pricing transparency | 0 | 0.08 | 0.0 | - |
Computed . Re-weight it by intent, or see the full method.
All data & sourcesshow ↓
Every value we hold, with its source, retrieval date, and confidence. This is the evidence behind the score: don't trust it, verify it.
Features
| Attribute | Value | Evidence |
|---|---|---|
| Capabilities | Architecture model: managed_enterprise_saas · Human in the loop hitl annotation ui: Yes · Llm as a judge prompt grading framework: Yes · Ci cd github actions pipeline blocking gates: Yes · Dashboard analytics for metric drift over time: Yes · Synthetic test dataset generation from documents: Yes | mediumsource · 2026-08-21 · 60% |
Integrations
| Attribute | Value | Evidence |
|---|---|---|
| Count | 1 | mediumsource · 2026-08-21 · 60% |