Product overview
What is Speechmatics?
Speechmatics is an enterprise speech-technology platform for developers building transcription, captioning, translation, analytics, meeting, accessibility, and voice-agent products. Its APIs cover real-time and batch speech-to-text, text-to-speech, and voice-agent workflows. Current pricing materials state support for more than 56 transcription languages and 69 translation pairs.
The platform provides language identification, speaker diarization, timestamps, custom dictionaries, formatting, audio-event handling, and optional summaries, chapters, sentiment, topics, and translation. Deployment choices range from managed multi-region cloud to private cloud, containers, virtual appliances, on-device options, and enterprise on-premises environments, depending on plan and capability.
How to Use Speechmatics
- Create an account and project, then issue an API key with the minimum required access.
- Choose real-time streaming or batch-file transcription and the appropriate accuracy model.
- Configure language identification, diarization, vocabulary, timestamps, formatting, and post-processing.
- Add translation, speech generation, or a voice-agent conversation layer when needed.
- Test noisy audio, accents, code-switching, latency, consent, and retention before production.
Core Features
- Speech-to-text APIs: Transcribes streaming audio and uploaded files.
- Multilingual coverage: Supports more than 56 languages plus multilingual and translation workflows.
- Voice-agent components: Supplies low-latency speech recognition and generation for conversational systems.
- Transcript enrichment: Includes speakers, timestamps, punctuation, numerals, audio events, and custom vocabulary.
- Flexible deployment: Offers SaaS and selected private or on-premises deployment models.
- Enterprise controls: Provides higher concurrency, custom models, support, and security-oriented deployment choices.
Use Cases
- Live captions and broadcast or event transcription.
- Contact-center analytics and searchable call records.
- Meeting assistants, medical scribes, and legal transcription.
- Multilingual voice agents, IVR, and customer-service automation.
Pricing
The Free plan includes 3,000 speech-to-text minutes per month split between real-time and batch, plus one million text-to-speech characters. Pro usage starts at $0.129 per hour for Batch Melia 1; other listed speech-to-text rates range from $0.24 to $0.43 per hour, with model-training and volume discounts. Enterprise pricing is custom and adds scale, private deployments, custom models, and prioritized support. Rates and allowances can change, so production estimates should use the current pricing page.
Frequently Asked Questions
Is Speechmatics only a transcription app?
No. It is primarily an API platform that also provides translation, text-to-speech, and voice-agent building blocks.
Can it run outside a shared SaaS environment?
Selected enterprise capabilities support private and on-premises deployment; availability varies by component.
Should transcripts be treated as exact records?
No. Accuracy varies with audio, speakers, vocabulary, and language. High-stakes records require review and source-audio retention policies.


