Documentation Navigation
Models & pricing / Models / Qwen3.8-2.4T-A95B-FP8
Qwen3.8-2.4T-A95B-FP8 Overview
This model's data has not been fully verified and benchmarks may be missing.
Qwen3.8-2.4T-A95B-FP8 is a chat-reasoning AI model created by Qwen, featuring 2.4T parameters and a context window of 128K tokens.
Qwen3.8-2.4T-A95B-FP8 is a 2.4T-parameter Mixture-of-Experts (MoE) text generation model developed by Qwen. Built on the Qwen3_5MoeForCausalLM architecture using transformers. Released on 2026-08-08 with 174 likes and 9,334 downloads on Hugging Face.
Model Lineage & Specification
qwen38-24t-a95b-fp8 is Qwen's primary release in the current family.
qwen38-24t-a95b-fp8Comparable models
| Feature | Qwen3.8-2.4T-A95B-FP8 | Solar Pro 4 | Qwen3.8 Max Preview | Nemotron 3 Ultra |
|---|---|---|---|---|
| Description | Qwen3.8-2.4T-A95B-FP8 is a 2.4T-parameter Mixture-of-Experts (MoE) text generati... | Solar Pro 4 is a large language model from Upstage built for enterprise workflow... | Preview Qwen flagship for million-token multimodal reasoning and long-horizon ag... | NVIDIA Nemotron 3 Ultra is a flagship open-weight frontier LLM family engineered... |
| API Identifier | qwen38-24t-a95b-fp8 | upstage-solar-pro-4 | alibaba-qwen3.8-max-preview | nvidia-nemotron-3-ultra |
| Parameters | 2.4T | — | — | 550B total / 55B active |
| Context Window | 128K tokens | 524K tokens | 1M tokens | 1M tokens |
| License / Type | open weights | api only | api only | open weights |
Qwen3.8-2.4T-A95B-FP8
Qwen3.8-2.4T-A95B-FP8 is a 2.4T-parameter Mixture-of-Experts (MoE) text generati...
qwen38-24t-a95b-fp8Solar Pro 4
Solar Pro 4 is a large language model from Upstage built for enterprise workflow...
upstage-solar-pro-4Qwen3.8 Max Preview
Preview Qwen flagship for million-token multimodal reasoning and long-horizon ag...
alibaba-qwen3.8-max-previewNemotron 3 Ultra
NVIDIA Nemotron 3 Ultra is a flagship open-weight frontier LLM family engineered...
nvidia-nemotron-3-ultraQwen3.8-2.4T-A95B-FP8 is a 2.4T-parameter Mixture-of-Experts (MoE) text generation model developed by Qwen. Built on the Qwen3_5MoeForCausalLM architecture using transformers. Released on 2026-08-08 with 174 likes and 9,334 downloads on Hugging Face. Powered by 2.4T parameters, Qwen3.8-2.4T-A95B-FP8 delivers specialized capabilities across chat reasoning with a native context window of 128K tokens. Built by Qwen, the architecture prioritizes low-latency throughput, dependable reasoning fidelity, and flexible deployment across enterprise APIs and local hardware environments.