INInkling
Inkling: Compact Thinking & Reasoning Engine
Inkling by Thinking Machines AI is a 7B parameter compact reasoning model fine-tuned for on-device agentic decision making, structured output generation, and low-latency logic tasks.
🧠 Key Features
- Built-in Chain of Thought: Emits internal reasoning traces (
<think>...</think>) before returning final outputs. - Structured JSON Mode: 100% schema enforcement for agentic tool calls.
- Fast Execution: Up to 120 tokens/sec on consumer GPU hardware.
📊 Benchmarks
| Benchmark | Score |
|---|---|
| MATH-500 | 78.2% |
| MBPP Code | 74.5% |
Key Features
Native MoE architecture with 975B total parameters and 41B active parameters
Pretrained on 45 trillion tokens including text, images, audio, and video
Support for 1 million tokens context window at full inference
You might also want to compare
Verified Sources
Tags
Model Specs
Parameters
975B (41B active)
Context Window
1000K tokens
License
Other/Custom
Deployment
Resources & Links
Lineage
Model Family
Part of the thinking-machines-inkling family
Only release in this line currently tracked.
Curator Notes
First open-weights model from Thinking Machines. Features a massive 975B MoE structure with sparse active routing.
Compare Specs
Compare parameters, context windows, modalities, and benchmark scores of this model side-by-side with others.
Compare Model