Documentation Navigation

Models & pricing / Models / Ling-3.0-flash

Ling-3.0-flash Overview

This model's data has not been fully verified and benchmarks may be missing.

Ling-3.0-flash is a chat-reasoning AI model created by inclusionAI, featuring 127.5B parameters and a context window of unknown.

vLLM is a fast and easy-to-use library for LLM inference and serving, originally developed in the Sky Computing Lab at UC Berkeley, featuring state-of-the-art serving throughput, efficient management of attention key and value memory, and support for various quantization methods and hardware platforms.

Model Lineage & Specification

ling-30-flash is inclusionAI's primary release in the current family.

DeveloperinclusionAI
API Identifierling-30-flash
Parameters127.5B
Context Windowunknown
LicenseUnknown

Comparable models

Ling-3.0-flash

vLLM is a fast and easy-to-use library for LLM inference and serving, originally...

API IDling-30-flash
Typeopen weights
Parameters127.5B
Contextunknown

Solar Pro 4

Solar Pro 4 is a large language model from Upstage built for enterprise workflow...

API IDupstage-solar-pro-4
Typeapi only
Parameters
Context524288

Qwen3.8 Max Preview

Preview Qwen flagship for million-token multimodal reasoning and long-horizon ag...

API IDalibaba-qwen3.8-max-preview
Typeapi only
Parameters
Context1M tokens

Nemotron 3 Ultra

NVIDIA Nemotron 3 Ultra is a flagship open-weight frontier LLM family engineered...

API IDnvidia-nemotron-3-ultra
Typeopen weights
Parameters550B total / 55B active
Context1M tokens

vLLM is a fast and easy-to-use library for LLM inference and serving, originally developed in the Sky Computing Lab at UC Berkeley, featuring state-of-the-art serving throughput, efficient management of attention key and value memory, and support for various quantization methods and hardware platforms.