Speechmatics
Multilingual speech-to-text and text-to-speech APIs deployable in cloud, on-premises, or on-device environments
- Also known as
- speechmatics
Available worldwide
What is Speechmatics?
Speech recognition and synthesis platform offering real-time and batch transcription, audio analysis, and text-to-speech across 55+ languages. Supports flexible deployment with privacy-first design and compliance-ready infrastructure.
Speechmatics pricing
Plans, per-tier features and add-ons, dated and linked to live pricing. Pricing changes often; always verify at source before you rely on it.
Usage-based from $0.24/hour; free tier with $100 credit; volume discounts available
Free
FreeFor developers and early exploration
- credit
- $100
- concurrent_real_time_sessions
- 2
- Speech-to-Text (55+ languages)
- Text-to-Speech (English)
- Multi-region cloud options
- 2 concurrent real-time sessions
Pro
For demanding projects and growing needs
- file_jobs_per_second
- 10
- concurrent_real_time_sessions
- 50
- Speech-to-Text (55+ languages)
- Text-to-Speech
- Real-time sessions (up to 50 concurrent)
- Speaker diarization
- Custom vocabulary
- Advanced punctuation and casing
- Precise timestamps
- Multi-channel support
- Online email support
- 20% volume discount over 500 hours/month
Enterprise
Contact salesCustom pricing with unlimited scale and flexible deployment options
- Unlimited concurrent real-time sessions
- No rate limits
- Custom models and voices
- Custom language development
- SaaS or on-premises or on-device deployment
- Privacy-first deployment options
- Dedicated Customer Success Manager
- Prioritized service and support
Add-ons
- Model Training Discount Program33% per Speech-to-Text usage
What Speechmatics does
The capabilities that matter for speech to text, normalised so it lines up with every alternative. “-” means we haven't confirmed it, not that it's missing.
- Real time streaming
- ✓
- Speaker diarization
- ✓
- Custom vocabulary
- ✓
- Fine tuning
- ✓
- Word timestamps
- ✓
- Punctuation formatting
- ✓
- Pii redaction
- -
- Vertical model
- ✓
- Offline on device
- ✓
- Open source
- -
Platform & deployment
Independently observed- CLI
- Web
- Cloud / SaaS
- Hybrid
- On-premise
- Self-hosted
Integrations (7)
Independently observed- Stenograph CATalyst VP
- Adobe Premiere
- Zapier
- Twilio
- Python SDK
- JavaScript SDK
- .NET SDK
Speechmatics alternatives
Other speech to text we track, ranked by the same independent score.
- AssemblyAIAPIs for converting audio to text and building voice-driven applications with real-time and recorded processingmedium · 60%
- whisper.cpplow · 24%
- DeepgramConversation-aware voice AI for transcription and synthesis at scalelow · 43%
- Amazon TranscribeAutomatic speech recognition service that converts spoken audio into written textlow · 34%
- faster-whisperlow · 28%
- Vosklow · 28%
Compare Speechmatics
Side by side against other speech to text, attribute by attribute, with a source on every value.
The Vioscale score: one lens on the evidence
Not user reviews and not a paid placement: a confidence-weighted blend of the independent signals below (adoption, activity, security posture, and more), which you can sort and re-weight yourself. Vendors can correct their listing but can never move their rank, and stars are weighted low as a vanity metric. It is one way to read the evidence for Speechmatics, not the verdict.
| Signal | Score | Weight | Contribution | Evidence |
|---|---|---|---|---|
| Pricing transparency | 100 | 0.08 | 8.4 | ✓ |
| Security posture | 85 | 0.07 | 6.2 | ✓ |
| Capabilities | 92 | 0.05 | 4.5 | ✓ |
| Price level | 80 | 0.05 | 4.2 | ✓ |
| Integrations | 26 | 0.04 | 1.1 | ✓ |
| Reliability | 0 | 0.07 | 0.0 | - |
Computed . Re-weight it by intent, or see the full method.
All data & sourcesshow ↓
Every value we hold, with its source, retrieval date, and confidence. This is the evidence behind the score: don't trust it, verify it.
Features
| Attribute | Value | Evidence |
|---|---|---|
| Capabilities | Fine tuning, Vertical model, Word timestamps, Custom vocabulary, Offline on device, Real time streaming, Speaker diarization, Punctuation formatting | mediumsource · 2026-08-25 · 60% |
Integrations
| Attribute | Value | Evidence |
|---|---|---|
| Count | 7 | mediumsource · 2026-08-25 · 60% |
Market
| Attribute | Value | Evidence |
|---|---|---|
| Availability | HqCountry: GB · PrimaryMarkets: … · AvailabilityScope: global · AvailableCountries: … · NotAvailableCountries: … | highsource · 2026-08-25 · 75% |