Useful Sensors
Useful Sensors
AudioOpen Weights (Apache 2.0) Verified Architecture & SpecsFree ($0 API Tokens)

Moonshine v2 Large STT

Ultra-lightweight, edge-native speech-to-text (STT) encoder-decoder model optimized for on-device real-time transcription and edge robotics with 5x lower compute latency than Whisper Large v3.

Technical Architecture & Execution Specifications
Architecture Overview

Moonshine v2 Large STT

Ultra-lightweight, edge-native speech-to-text (STT) encoder-decoder model optimized for on-device real-time transcription and edge robotics with 5x lower compute latency than Whisper Large v3.

Memory Math Breakdown
  • • FP16 Weights = 0.5B × 2B = 0.96 GB
  • • INT4 Weights = 0.5B × 0.55B = 0.26 GB
  • • KV Cache (65536 ctx, FP16) ≈ 8.00 GB
  • • Activation Buffer = ~20% overhead
Supported Modalities
audiotext

Hardware & Execution ParametersAudio

Total Parameter Count480M
Active Parameters (MoE)480M per token
Context Window Capacity65,536 tokens
Model Weights Footprint1.0 GB (FP16) / 0.3 GB (INT4)
Distribution LicenseOpen Weights (Apache 2.0)
Standard API Pricing (1M Tokens)$0 in / $0 out
Model Heritage & Evolutionary Lineage
Moonshine v2

Genealogical Graph & Evolutionary Provenance

Tracing foundational base architecture ancestry, architectural successors, scale siblings, and reasoning distillation derivatives.

Independent Foundation Checkpoint

Moonshine v2 Large STT operates as an autonomous foundation model architecture without direct precursor derivatives in this catalog.

Root Architecture Node
AI Model Architecture & Intelligence

Architecture Engineering & Capability Deep-Dive

An objective architectural evaluation of Moonshine v2 Large STT by Useful Sensors, analyzing underlying compute dynamics, memory constraints, and deployment economics.

Topology & Attention Mechanics

A robust autoregressive transformer utilizing standard attention patterns for predictable and coherent token generation.

Architecture Type:Advanced Transformer

Evaluation Profile & Reasoning

Exhibits frontier-tier behavior in reasoning and coding.

Domain Specialty:Audio

LLM Hardware Sizing & Serving

For open deployments via vLLM/SGLang, quantization (INT4/AWQ) is heavily recommended to fit dense memory constraints, or multi-GPU pipeline parallelism for full FP16.

KV Cache Mgmt: PagedAttention / FlashAttention-3
Hosting Type:Open Weights (Apache 2.0)

Inference Economics & Workflows

Well-suited for enterprise pipelines where capability is balanced against per-million token costs.

Enterprise Fit:Production Ready
Production Trade-Offs & Capability Balance

Architectural Strengths vs. Considerations

An objective balance sheet analyzing the operational advantages and production constraints of deploying Moonshine v2 Large STT.

Key Architectural Strengths

  • Ultra cost-effective inference at $0/1M input tokens enables high-frequency agent loops.
  • Demonstrated RTF (Real-Time Factor) evaluation score of 0.08% in verified benchmarks.

Operational Considerations

  • Non-deterministic reasoning chains require schema validation in safety-critical deployments.
LLM Hardware Sizing & Runtime Compatibility

Inference Runtimes & Hardware Sizing

Deployment targets, inference engines, and memory requirements for Moonshine v2 Large STT.

Recommended Hardware Profile:
1x RTX 4070 / 4080 (16GB) or Mac M-Series (16GB Unified)
Consumer GPU (< 16 GB VRAM)
Est. 1.2 GB (FP16) / 0.3 GB (INT4)
Production Serving Recipes
vLLM Production
python3 -m vllm.entrypoints.openai.api_server --model Moonshine v2 Large STT --tensor-parallel-size 1 --gpu-memory-utilization 0.95 --max-model-len 4096 --enable-chunked-prefill
Ollama / llama.cpp
ollama run moonshine v2 large stt
SGLang Structured
python3 -m sglang.launch_server --model-path Moonshine v2 Large STT --tp 1 --trust-remote-code
TGI Serving
text-generation-launcher --model-id Moonshine v2 Large STT --num-shard 1 --max-batch-prefill-tokens 32000
Supported Inference Engines
vLLM

High-throughput PagedAttention server

Supported
Ollama

One-click CLI & local desktop serving

Supported
SGLang

Fast multi-turn structured decoding

Supported
TGI

Text Generation Inference

Supported
Llama.cpp

GGUF CPU/Apple Silicon execution

Supported
Precision & Quantization Formats
BF16 / FP16Full Precision

~1.2 GB VRAM required

FP8 (E4M3)Native FP8

~0.6 GB VRAM (Hopper speedup)

AWQ / GPTQ (4-bit)Activation-Aware

~0.3 GB VRAM

GGUF (Q4_K_M)Quantized Binary

CPU RAM / Apple Silicon optimized

LLM Benchmark Database & Performance Metrics

3 Tested

Standardized evaluation results across reasoning, agentic coding, computer use, and alignment.

Flagship Headline MetricsIndustry SOTA Standard
General / Other

CommonVoice Multilingual WER %

4.15%
0%100%
General / Other

LibriSpeech Clean WER %

1.82%
0%100%
General / Other

RTF (Real-Time Factor)

0.08%
0%100%
RTF (Real-Time Factor)
0.08%
LibriSpeech Clean WER %
1.82%
CommonVoice Multilingual WER %
4.15%
Commercial Rates & Inference Costs

API & Deployment Pricing

Open-weights model available for local and private cloud deployment. Compute costs depend on the target GPU hardware instance.

Usage TierRate / Unit
Prompt / Input Tokens$0 / 1M tokens
Completion / Output Tokens$0 / 1M tokens
Inference Cost & ROI Engine
Market Cloud Rate
Prompt / Input Volume:50M Tokens / mo
1M500M1,000M
Generated / Output Volume:10M Tokens / mo
1M250M500M
Estimated Monthly Spend
$160.00/ mo
Input (50M @ $2.00/1M):$100.00
Output (10M @ $6.00/1M):$60.00
Cloud GPU Breakeven Ratio
RunPod RTX 4090 ($316/mo)0.51x spend
Lambda 1x H100 ($1,800/mo)0.09x spend
💡 Open-weights model. You can self-host for $0 token API charge or consume via managed serverless endpoints at the rates shown above.
Similar Frontier Models & Alternatives
Explore All Comparisons

Comparable Foundation Architectures

Alternative models in the Audio class with similar capabilities, context windows, or deployment profiles.

MetaAudio

Muse Voice Transcribe

Context:Standard ctx
Parameters:Undisclosed
Input Rate:Free / Self-Host
OpenAIAudio

GPT-Realtime-Translate

Context:16k ctx
Parameters:Proprietary
Input Rate:Free / Self-Host
OpenAIAudio

GPT-Live-Transcribe

Context:Standard ctx
Parameters:Proprietary
Input Rate:Free / Self-Host
Integration & Deployment
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ.get("USEFUL_SENSORS_KEY", "EMPTY"),
    base_url="http://localhost:8000/v1"
)

response = client.chat.completions.create(
    model="moonshine-v2-speech-stt-2026",
    messages=[{"role": "user", "content": "Explain quantum superposition in 2 sentences."}]
)
print(response.choices[0].message.content)
Frequently Asked Questions

Frequently Asked Questions about Moonshine v2 Large STT

Essential facts, architectural specs, hardware constraints, and pricing answers for Moonshine v2 Large STT.

To run Moonshine v2 Large STT (480M) locally, you generally need Depends on quantization. We recommend using quantized GGUF/AWQ formats with Ollama or vLLM to optimize memory footprint.

Primary Sources & Access Repositories

All technical specifications, parameter distributions, context architectures, and benchmark evaluations for Moonshine v2 Large STT are audited against primary source release documentation, research whitepapers, and verified vendor API endpoints.