Galileo

AI evaluation and observability platform for testing, monitoring, and improving AI systems

Also known as
galileo

What is Galileo?

A platform for building and evaluating AI systems that synthesizes test datasets from multiple sources, compresses expensive LLM evaluators into efficient models, and provides production monitoring with real-time observability.

Independently observed

What Galileo does

The capabilities that matter for ai evals testing, normalised so it lines up with every alternative. “-” means we haven't confirmed it, not that it's missing.

Capabilities
Architecture model
Managed enterprise saas
LLM as a judge prompt grading framework
Specialized rag metrics faithfulness context relevance
-
Deterministic regex and json schema assertions
-
Synthetic test dataset generation from documents
Ci cd github actions pipeline blocking gates
Red teaming and adversarial vulnerability scanning
-
Multi model side by side ab regression testing
-
Human in the loop hitl annotation UI
Dashboard analytics for metric drift over time
SOC2 type ii
-
Mit or apache permissive oss license
-
Pricing model
-
Independently observed

Integrations (1)

Independently observed
  • NVIDIA NeMo

Galileo alternatives

Other ai evals testing we track, ranked by the same independent score.

All Galileo alternatives, ranked →

Independent · unbought · dated

The Vioscale score: one lens on the evidence

Not user reviews and not a paid placement: a confidence-weighted blend of the independent signals below (adoption, activity, security posture, and more), which you can sort and re-weight yourself. Vendors can correct their listing but can never move their rank, and stars are weighted low as a vanity metric. It is one way to read the evidence for Galileo, not the verdict.

Balanced composite 45 / 100
low · 8%updating
Signal contributions to the composite score
SignalScoreWeightContributionEvidence
Capabilities750.053.6
Integrations90.040.4
Price level00.050.0-
Reliability00.070.0-
Security posture00.070.0-
Pricing transparency00.080.0-

Computed . Re-weight it by intent, or see the full method.

All data & sourcesshow ↓

Every value we hold, with its source, retrieval date, and confidence. This is the evidence behind the score: don't trust it, verify it.

Features

AttributeValueEvidence
CapabilitiesArchitecture model: managed_enterprise_saas · Human in the loop hitl annotation ui: Yes · Llm as a judge prompt grading framework: Yes · Ci cd github actions pipeline blocking gates: Yes · Dashboard analytics for metric drift over time: Yes · Synthetic test dataset generation from documents: Yesmediumsource · 2026-08-21 · 60%

Integrations

AttributeValueEvidence
Count1mediumsource · 2026-08-21 · 60%