Method

How the vioscaleAI score works

The composite is a confidence-weighted blend of independent signals, not user reviews, not a vendor's self-description. Each signal is normalised to 0-100, multiplied by a category-specific weight, and combined. The score you see is the sum of contributions, discounted by how much evidence we actually have.

This mechanism is the same for any software category. What changes per category is which signals count and how much they are weighted, which is why the score is neutral across very different kinds of software. A category's weighting is chosen to match how that market is actually evaluated: developer tooling leans on adoption and development activity, security software on audited posture, infrastructure on reliability, and business verticals on price, features and integration breadth. The signals below are the developer-tooling set, shown because it is the most familiar.

Evidence changes the score, not just the confidence

A composite built from three signals is not comparable to one built from twelve, yet a ranking places them side by side. So the published score is the measured average pulled toward a neutral 50 in proportion to how much of the category's applicable evidence we actually hold. A fully measured product is unaffected. A barely measured one sits near the middle until we know more.

This is deliberately symmetric, not a penalty. A thinly measured product that happened to look excellent is pulled down; one that happened to look poor is pulled up by exactly the same rule. We are not punishing a product for gaps in our own crawling; we are declining to claim more precision than the evidence supports, in either direction.

The signals we score today (developer tooling)

These are the signals for developer tooling, shown as a worked example. Weights are set per category, so security posture matters more for CI/CD than for a web framework, and stars are always weighted low.

SignalTypical weightWhy it counts
Package downloadshighThe strongest independent adoption signal: real usage, hard to fake at scale.
Dependent projectsmediumHow many public projects depend on it: cross-ecosystem downstream adoption that counts real usage, so it is hard to fake.
Developer Q&A activitymediumPractitioner mind-share and the size of the help ecosystem.
Development activitymediumCommit and contributor velocity: is the project alive and maintained?
Release cadencemediumHow regularly the project actually ships. Momentum, not promises.
Security posturemedium-highSelf-reported SOC 2 / ISO 27001, disclosure policy, licence; weighted higher where it matters (CI/CD, ORMs).
Security scoremediumAn independently measured security-practices score (branch protection, code review, signed releases, CI checks) that complements the self-reported certifications above.
Docs qualitylow-mediumHow usable the software is once you adopt it.
Starslow (vanity)Deliberately weighted low. Stars are a popularity artefact, easily gamed and weakly correlated with fitness.

Other verticals score a different set. The CRM, project management, and analytics categories, for example, weigh signals like pricing transparency, compliance certifications (SOC 2, ISO 27001, GDPR, HIPAA), reliability (a public status page and SLA), and integration breadth, not raw download counts. The scoring machinery, provenance, confidence, and freshness is identical; only the signal set and weights change.

Confidence bands

We never express false precision. Each fact and each score carries a confidence in one of four bands. Missing evidence lowers confidence rather than silently assuming a value.

high · 85%medium · 60%low · 30%unverified
  • high: ≥ 75%
  • medium: 50-74%
  • low: 1-49%
  • unverified: no evidence observed

When we refuse to rank

A numbered leaderboard is a claim about every pair in it, so it has to clear the same bar we apply to a head-to-head comparison. Where the best-evidenced product in a category still falls below that bar, we do not publish a ranking at all: the products are listed with their evidence and the reason is stated on the page. You will see this on categories we have only begun to cover. It is not an error, and it is not hidden.

The same rule governs a comparison: where two products are within a few points, or the leader's evidence is too thin, we say there is no clear winner rather than manufacturing one.

Intent-conditioned scoring

A score is not a single truth: it is one weighting of the evidence. So the facts and signals stay fixed and provenanced, and only the weight vector varies. The balanced profile is the default and the canonical number we and anyone else should cite. It never changes per query. On top of that, an agent can ask for a re-weighted ranking, by a named intent or an explicit weight vector, and the underlying numbers are identical: only the emphasis moves.

Every profile is public (there are no secretly tuned weights), so no vendor can buy a favourable one. A single-signal cap of 0.5 means one gameable signal can never dominate a popular intent, and the resolved, normalised weights ship in the response so a consumer can verify the ranking. The same profiles are available to agents over the API and MCP, and to you on any category page; the balanced score remains the canonical one to cite.

An intent asks a specific question, so a product we hold no evidence on for that question is not ranked against it. Ask for the most secure tools and a product whose security posture we have never observed is listed separately, marked as such, rather than being quietly slotted in on its other merits. This means an intent often returns fewer products than the balanced view. That is the honest answer: re-weighting evidence we do not have would move nothing while looking like it had.

ProfileWhat it emphasises
balancedThe default, citeable composite. Category base weights, unchanged.
most-adoptedReal-world usage: package downloads and developer Q&A activity.
most-activeDevelopment velocity: recent commit activity and release cadence.
most-secureSecurity and compliance posture: SOC 2, ISO 27001, disclosure policy.
best-valueQuality per cost. Leans on adoption as a proxy until price data lands in paid categories.

Agents can also skip the named profiles and pass explicit weights, e.g. weights=security_posture:0.4,package_downloads:0.3,release_cadence:0.3. Every profile, its multipliers, and the signal glossary are published, machine-readable, at /api/v1/intents. See the AI & agents page for the query params and examples.

Freshness

A fact's trust decays with age until it is re-crawled. Every value records when we last observed it, and the serving surface exposes that date so a reader, human or machine, can judge staleness for itself. Facts older than the freshness window are re-fetched on a schedule.

A worked example

95
95 / 100 · medium confidence

Here is exactly how that number is built: signal by signal, weight by weight, with contribution shown. Nothing is hidden:

Signal contributions to the composite score
SignalScoreWeightContributionEvidence
Reliability980.1615.3
Security posture1000.1211.5
Pricing transparency1000.066.3
Capabilities1000.055.1
Integrations1000.054.9
Price level500.042.1

Machine: /api/v1/signals/oneuptime

Indexed uniformly, owned by no vendor

Independence is a promise you should be able to check, so here is exactly how vendors fit in, in plain terms.

  • Uniform and free. Every product is indexed and scored by the same algorithm, for free, whether or not its vendor has claimed it. A claimed profile and an unclaimed one are measured identically.
  • Vendors control their own profile, not the market. A vendor can claim their profile to correct, refine, and keep it current. That governs how their own product is presented, never where it sits in a ranking or a comparison.
  • Vendor content is labelled and kept separate. Anything a vendor supplies is clearly marked as provided by the vendor and kept distinct from the facts we indexed independently, so you can always tell a self-claim from an observation and weigh them accordingly.
  • None of it moves rank, score, or position. Claiming and editing change only a vendor's own presentation. Signals stay algorithm-computed and vendor-locked, so no amount of vendor control moves a rank, a score, or a place in a comparison.

Provenance

Provenance is the moat. Every value vioscaleAI publishes is wrapped so it can be cited: what the value is, where it came from, when we last saw it, and how confident we are.

A provenanced fact carries
  • attribute: the namespaced key, e.g. pricing.model
  • value: the observed value
  • source: the public URL where we observed it
  • retrievedAt: ISO-8601 timestamp of last observation
  • confidence: 0-1, surfaced as a band
  • sourceType: who asserted it (crawler, vendor, system)

This evidence travels with the data across every format: visible on the HTML page, embedded in JSON-LD as a vioscale:provenance block, and inline in the Markdown twin. See how to consume it on the AI & agents page.