AITTS-1
TTS-1: OpenAI's Text-to-Speech Model
Model Overview
TTS-1 is OpenAI's text-to-speech API model, released in November 2023. It converts text input into realistic, natural-sounding speech audio in 6 distinct voices. Optimized for real-time generation with low latency, TTS-1 supports up to 4096 characters per request and outputs audio in multiple formats (MP3, Opus, AAC, FLAC, WAV, PCM). A higher-quality variant, TTS-1-HD, offers improved audio fidelity at higher latency for non-real-time applications.
✨ Key Features
| Feature | Description |
|---|---|
| 6 Voices | Choose from Alloy, Echo, Fable, Onyx, Nova, and Shimmer voices |
| Low Latency | Optimized for real-time streaming audio generation |
| Multiple Formats | Supports MP3, Opus, AAC, FLAC, WAV, and PCM output |
| Multilingual | Generates speech in multiple languages following the text input |
| Two Tiers | TTS-1 (fast) and TTS-1-HD (higher quality) for different use cases |
🔗 Resources
| Resource | Link |
|---|---|
| API Docs | platform.openai.com/docs/guides/text-to-speech |
📜 License & Access
Proprietary — Available via OpenAI API.
You might also want to compare
Verified Sources
Model Specs
Parameters
Unknown
Context Window
Unknown
License
Proprietary
Deployment
Cost Tiers
Resources & Links
Lineage
Model Family
Part of the openai-tts family
Only release in this line currently tracked.
Curator Notes
Bulk imported from OpenAI developer docs.
Compare Specs
Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.
Compare Model