Back to Academic/Research
Research PreviewAudio & SpeechtextaudioUpdated June 15, 2026

Dots TTS (dots.tts): Continuous Space Text-to-Speech Foundation Model

Dots TTS (dots.tts) is a 2-billion parameter, fully continuous, end-to-end autoregressive text-to-speech (TTS) foundation model developed by RedNote HiLab (the AI research division of Xiaohongshu).

Traditional TTS models often rely on discrete speech codec tokens (such as EnCodec or SoundStream), which can introduce acoustic information loss similar to image pixelation. In contrast, Dots TTS operates entirely within a continuous latent space, using a 48 kHz AudioVAE combined with an autoregressive flow-matching head over a Qwen2.5-1.5B LLM backbone. This design achieves exceptionally high acoustic fidelity, natural intonation, and superior zero-shot speaker cloning.


🔬 Key Features & Architecture

  • Fully Continuous Latent Representation: Operates over continuous audio embeddings via a high-resolution 48 kHz AudioVAE, eliminating discrete tokenization artifacts.
  • Qwen2.5 Backbone + Flow-Matching Acoustic Head: Combines semantic reasoning capabilities of a Qwen2.5 LLM backbone with autoregressive flow-matching for detailed audio synthesis.
  • Self-Corrective Alignment (SCA): Incorporated into the dots.tts-soar model variant to minimize alignment drift and enhance speaker identity preservation.
  • MeanFlow Low-Latency Distillation: Powers dots.tts-mf, enabling few-step inference for real-time speech generation with first-packet latency as low as 54 ms.
  • Multilingual Zero-Shot Voice Cloning: Supports 24 languages with prompt-based zero-shot voice cloning from short audio reference clips.
bash
Text Input + Audio Reference Prompt ───► Qwen2.5-1.5B LLM ───► Flow-Matching Acoustic Head ───► 48kHz AudioVAE Decoder ───► High-Fidelity Audio

📊 Model Variants & Checkpoints

Specification Table
Checkpoint NamePurpose & Performance
dots.tts-soarRecommended Default. Refined via Self-Corrective Alignment for highest fidelity and speaker similarity.
dots.tts-mfReal-Time Streaming. Distilled with MeanFlow for ultra-fast generation (54ms first-packet latency).
dots.tts-baseBase Foundation Model. Pretrained 2B checkpoint for fine-tuning.

🚀 Quickstart & Code Usage

bash
pip install dots.tts
python
from dots_tts.runtime import DotsTtsRuntime
import soundfile as sf

# Load pre-trained model (e.g. dots.tts-soar)
runtime = DotsTtsRuntime.from_pretrained("rednote-hilab/dots.tts-soar")

# Generate speech from text and prompt audio
audio = runtime.synthesize(
    text="Welcome to the future of continuous space text-to-speech synthesis.",
    prompt_audio="reference_voice.wav",
    prompt_text="Transcript of reference voice sample."
)

# Export audio file at 48kHz
sf.write("output.wav", audio, samplerate=48000)
print("Speech synthesis complete: output.wav")

🔗 Paper & Resources

Key Features

Fully Continuous Latent Space: Replaces discrete codec tokens with a continuous 48 kHz AudioVAE representation

Feature 01

Autoregressive Flow-Matching Architecture: Integrates a Qwen2.5-1.5B LLM backbone with an autoregressive flow-matching acoustic head

Feature 02

Self-Corrective Alignment (SCA): Refined alignment strategy in dots.tts-soar checkpoint for superior zero-shot speaker similarity

Feature 03

MeanFlow Low-Latency Distillation: Enables few-step generation in dots.tts-mf with first-packet latency as low as 54ms

Feature 04

Multilingual Zero-Shot Voice Cloning: Native support for 24 languages with prompt-based acoustic adaptation

Feature 05

You might also want to compare

Verified Sources

Tags

research-previewaudio-speechtext-to-speechzero-shot-voice-cloningflow-matchingaudio-vae

Model Specs

research-preview

Parameters

2B

Context Window

undisclosed

License

Apache-2.0

Deployment

self-hostable

Resources & Links

Curator Notes

Verified academic research paper and open-source project from RedNote HiLab (Xiaohongshu). Paper available on arXiv (2606.07080), GitHub repository under Apache-2.0 license, and model weights on Hugging Face.

Compare Specs

Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.

Compare Model