Documentation Navigation

Models & pricing / Models / Qwen3.8-2.4T-A95B-FP8

Qwen3.8-2.4T-A95B-FP8 Overview

This model's data has not been fully verified and benchmarks may be missing.

Qwen3.8-2.4T-A95B-FP8 is a chat-reasoning AI model created by Qwen, featuring 2.4T parameters and a context window of 128K tokens.

Qwen3.8-2.4T-A95B-FP8 is a 2.4T-parameter Mixture-of-Experts (MoE) text generation model developed by Qwen. Built on the Qwen3_5MoeForCausalLM architecture using transformers. Released on 2026-08-08 with 174 likes and 9,334 downloads on Hugging Face.

Model Lineage & Specification

qwen38-24t-a95b-fp8 is Qwen's primary release in the current family.

DeveloperQwen
API Identifierqwen38-24t-a95b-fp8
Parameters2.4T
Context Window128K tokens
LicenseApache-2.0

Comparable models

Qwen3.8-2.4T-A95B-FP8

Qwen3.8-2.4T-A95B-FP8 is a 2.4T-parameter Mixture-of-Experts (MoE) text generati...

API IDqwen38-24t-a95b-fp8
Typeopen weights
Parameters2.4T
Context128K tokens

Solar Pro 4

Solar Pro 4 is a large language model from Upstage built for enterprise workflow...

API IDupstage-solar-pro-4
Typeapi only
Parameters
Context524K tokens

Qwen3.8 Max Preview

Preview Qwen flagship for million-token multimodal reasoning and long-horizon ag...

API IDalibaba-qwen3.8-max-preview
Typeapi only
Parameters
Context1M tokens

Nemotron 3 Ultra

NVIDIA Nemotron 3 Ultra is a flagship open-weight frontier LLM family engineered...

API IDnvidia-nemotron-3-ultra
Typeopen weights
Parameters550B total / 55B active
Context1M tokens

Qwen3.8-2.4T-A95B-FP8 is a 2.4T-parameter Mixture-of-Experts (MoE) text generation model developed by Qwen. Built on the Qwen3_5MoeForCausalLM architecture using transformers. Released on 2026-08-08 with 174 likes and 9,334 downloads on Hugging Face. Powered by 2.4T parameters, Qwen3.8-2.4T-A95B-FP8 delivers specialized capabilities across chat reasoning with a native context window of 128K tokens. Built by Qwen, the architecture prioritizes low-latency throughput, dependable reasoning fidelity, and flexible deployment across enterprise APIs and local hardware environments.