Back to Qwen
Open WeightsTranslationaudiotextvisionUpdated May 15, 2026

Qwen Live Translate: Real-Time Multimodal Speech Translation

Model Overview

Qwen Live Translate (Qwen3.5-LiveTranslate-Flash) is Alibaba's real-time multimodal simultaneous speech-to-speech and speech-to-text translation model family built on the Qwen-Omni architecture by the Qwen Team (Alibaba Tongyi Lab).

It fuses audio, text, and visual input (such as speaker lip movements and facial context) to provide low-latency simultaneous interpretation, cross-lingual voice cloning, and domain-specific terminology control for live streaming and international conferences.


Key Features

  • Multimodal Audio-Visual Understanding: Integrates audio, text, and visual cues (lip movements, on-screen text) to resolve translation ambiguities in noisy environments.
  • Ultra-Low Latency Streaming: Utilizes "Readable Unit" chunk-based streaming to achieve average end-to-end speech-to-speech latency as low as 2.8 seconds.
  • Multilingual Support: Supports 60 languages for audio input and text translation, and 29 languages for speech audio synthesis.
  • Cross-Lingual Zero-Shot Voice Cloning: Automatically replicates the speaker's vocal characteristics (timbre and voice identity) in real-time.
  • Hotword Customization: Configures custom hotwords and domain terms to guarantee high accuracy in technical contexts.

Verified Project Links


Performance & Benchmarks

  • FLEURS & CoVoST 2: Achieves state-of-the-art BLEU/COMET scores across major speech translation directions.
  • End-to-End Latency: 2.8 seconds average end-to-end latency.

Key Features

Multimodal Audio-Visual Understanding: Integrates audio, text, and visual cues (lip movements, on-screen text) to resolve translation ambiguities

Feature 01

Ultra-Low Latency Streaming: Utilizes 'Readable Unit' chunk-based streaming to achieve end-to-end speech-to-speech latency down to 2.8s

Feature 02

Multilingual Support: Supports 60 languages for text translation and 29 languages for speech audio synthesis

Feature 03

Cross-Lingual Zero-Shot Voice Cloning: Automatically replicates speaker vocal characteristics in the translated output

Feature 04

Hotword Customization: Allows configuring custom hotwords and technical terminology for high accuracy

Feature 05

Verified Sources

Tags

translationspeech-to-speechvoice-cloningalibabaqwen

Model Specs

open-weights

Parameters

Undisclosed

Context Window

undisclosed

License

Apache-2.0

Deployment

api-onlyself-hostable

Resources & Links

Lineage

Curator Notes

Verified paper arXiv:2604.15804 and open-weights release from Alibaba Qwen Team.

Compare Specs

Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.

Compare Model