Back to Academic/Research
Open WeightsAudio & SpeechtextaudioUpdated June 1, 2026

WavTTS: Raw Waveform Text-to-Speech

Model Overview

WavTTS is an end-to-end zero-shot Text-to-Speech (TTS) synthesis framework developed by researchers at Shanghai Jiao Tong University (SJTU) and ByteDance Seed (Wenxi Chen, Dongya Jia, et al.). It generates speech directly in raw waveform space using flow matching and Diffusion Transformers (DiT), bypassing intermediate mel-spectrogram or neural codec token representations.


Key Features

  • Direct Raw Waveform Generation: Models continuous time-domain raw audio waveforms directly using Flow Matching and Diffusion Transformers (DiT).
  • Non-Overlapping Patchification: Employs a temporal patchification strategy to efficiently process high-dimensional raw audio signals.
  • Multi-Scale Mel Supervision: Integrates multi-scale mel-spectrogram loss during training to provide perceptual guidance without restricting generation at inference.
  • Zero-Shot Voice Cloning: Synthesizes natural, high-fidelity speech matching a target speaker's voice using only a short reference audio prompt.
  • Built on F5-TTS Framework: Extends the F5-TTS flow-matching framework for robust zero-shot audio synthesis.

Verified Project Links


Performance & Benchmarks

  • LibriSpeech-PC Word Error Rate (WER): 1.50% – 2.02%
  • UTMOS (Speech Naturalness): 4.36 (outperforms direct raw waveform baselines and matches top latent-space TTS models).

Key Features

Direct raw waveform modeling via flow matching and Diffusion Transformers (DiT)

Feature 01

Non-overlapping waveform patchification for long audio sequence processing

Feature 02

Multi-scale mel-spectrogram supervision for perceptual training guidance

Feature 03

Zero-shot voice cloning with short reference audio prompts

Feature 04

Built upon the F5-TTS architecture framework

Feature 05

You might also want to compare

Verified Sources

Tags

research-previewttstext-to-speechzero-shotaudio-generationsjtubytedance

Model Specs

open-weights

Parameters

Undisclosed

Context Window

undisclosed

License

MIT

Deployment

self-hostable

Resources & Links

Curator Notes

Verified paper arXiv:2606.03455 and open-source project from SJTU and ByteDance Seed.

Compare Specs

Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.

Compare Model