Back to thinking-machines-inkling
Open WeightsMultimodaltextimageaudioUpdated July 15, 2026
Some details or benchmark scores on this page are self-reported by developers and unconfirmed.

Inkling: Compact Thinking & Reasoning Engine

Inkling by Thinking Machines AI is a 7B parameter compact reasoning model fine-tuned for on-device agentic decision making, structured output generation, and low-latency logic tasks.


🧠 Key Features

  • Built-in Chain of Thought: Emits internal reasoning traces (<think>...</think>) before returning final outputs.
  • Structured JSON Mode: 100% schema enforcement for agentic tool calls.
  • Fast Execution: Up to 120 tokens/sec on consumer GPU hardware.

📊 Benchmarks

Specification Table
BenchmarkScore
MATH-50078.2%
MBPP Code74.5%

Key Features

Native MoE architecture with 975B total parameters and 41B active parameters

Feature 01

Pretrained on 45 trillion tokens including text, images, audio, and video

Feature 02

Support for 1 million tokens context window at full inference

Feature 03

You might also want to compare

Verified Sources

Tags

moemultimodaltinker

Model Specs

open-weights

Parameters

975B (41B active)

Context Window

1000K tokens

License

Other/Custom

Deployment

self-hostable

Resources & Links

Lineage

Model Family

Part of the thinking-machines-inkling family

Only release in this line currently tracked.

Curator Notes

First open-weights model from Thinking Machines. Features a massive 975B MoE structure with sparse active routing.

Compare Specs

Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.

Compare Model