Back to bonsai
PrismML /

BO
Bonsai 27B

Open WeightsAgentictextimageUpdated July 14, 2026
Some details or benchmark scores on this page are self-reported by developers and unconfirmed.

Bonsai 27B: Ultra-Efficient Sparse Frontier LLM

Bonsai 27B is PrismML's 27-billion parameter sparse mixture-of-experts model engineered to deliver frontier-class reasoning with 4× lower memory and inference footprint.


🌲 Architecture & MoE Efficiency

  • Total Parameters: 27B parameters.
  • Active Parameters: 4.1B active parameters per token.
  • Context Window: 128k tokens with flash-attention support.

📊 Benchmark Results

Specification Table
BenchmarkTask DomainBonsai 27BDense 70B Baseline
MMLUGeneral Knowledge82.4%81.9%
GSM8KMath Reasoning88.6%86.2%
HumanEvalCode Generation84.1%81.5%

Key Features

Ternary (1.71 effective bits) and 1-bit (1.125 effective bits) variants with no higher-precision escape hatches

Feature 01

Multimodal with a compact 4-bit vision tower

Feature 02

Supports multi-step reasoning, structured tool calls, and computer-use agentic loops

Feature 03

Native support on Apple devices via MLX and NVIDIA GPUs via CUDA

Feature 04

Supports speculative decoding

Feature 05

You might also want to compare

Verified Sources

Tags

on-deviceagentic1-bitternarymultimodalapple-mlxcuda

Model Specs

open-weights

Parameters

27B

Context Window

262K tokens

License

Apache 2.0

Deployment

on-deviceself-hostable

Resources & Links

Lineage

Model Family

Part of the bonsai family

Only release in this line currently tracked.

Compare Specs

Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.

Compare Model