Cartesia released Sonic-3.6, the successor to Sonic-3.5 that appeared roughly three months earlier. The update targets naturalness in generated speech and is validated by independent evaluation. On the Artificial Analysis speech leaderboards, Sonic-3.6 ranks first with 1,283 Elo on the Provider Voice board and 1,123 Elo on the Controlled Voice board, where all models are synthesized using the same eight reference voices to isolate the synthesis engine.
Unlike transformer‑based TTS systems, Sonic-3.6 is built on state space models. Cartesia reports a sub‑90 ms time‑to‑first‑audio and a 100 ms transcript latency for its companion Ink‑2 speech‑to‑text model. The model is offered only as a beta hosted API; no open weights or self‑hosted option are provided. Pricing is set at $49 per million characters, which is half the cost of ElevenLabs Eleven v3 and above Speechify Simba 3.2’s $10 per million characters.
- Architecture: State Space Model (non‑transformer)
- Latency: <90 ms time‑to‑first‑audio; 100 ms transcript latency (Ink‑2)
- Leaderboard Elo: Provider Voice 1,283; Controlled Voice 1,123
- Pricing: $49 per 1 M characters
- Availability: Beta hosted API only
- Licensing: Closed commercial; no open weights, no Hugging Face repo
Why this matters
The Controlled Voice board result shows that, when the voice catalog is held constant, Sonic-3.6’s synthesis engine outperforms competing models—a fact derived from the leaderboard. The vendor‑stated sub‑90 ms time‑to‑first‑audio and $49/M‑char pricing are also factual claims from Cartesia. Inferring from these points, the state‑space architecture appears able to deliver high perceptual quality while meeting low‑latency targets, which could benefit real‑time agent applications where turn‑taking speed is critical. However, because the latency figures are measured in isolation and not as end‑to‑end round‑trips, actual user‑perceived delay may vary depending on network conditions and client‑side processing.
