Back to Newsroom

Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks

By Modelverse Editorial·August 4, 2026·2 min read
Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks

Cursor Research has unveiled and open-sourced Mixture-of-Kittens (MoK), a groundbreaking mixture-of-experts (MoE) training megakernel that powers their Composer models. This release is significant for its reported performance, achieving up to 2.37 times higher throughput compared to existing public baselines, a testament to its efficiency in handling the complex demands of MoE architectures. MoK is now available on GitHub under the Apache-2.0 license, marking a crucial contribution to the AI community.

At its core, MoK revolutionizes MoE training by fusing every communication and computation step into a single, deterministic kernel. This innovative approach is specifically engineered for high-end NVIDIA Blackwell SM100 or SM103 GPUs, requiring GB200 or GB300 NVL72 racks. By leveraging Blackwell's Cluster Launch Control and PyTorch symmetric memory for inter-GPU buffers, MoK aggressively minimizes CPU-GPU synchronization, a critical bottleneck in previous MoE training setups. This design choice directly addresses the challenge of communication overhead, which can consume over half of the end-to-end training time for MoE layers.

MoK's release is particularly impactful for developers and researchers operating at the cutting edge of AI. While its hardware requirements limit adoption to organizations with substantial NVL72 capacity—such as frontier labs, GPU neoclouds, and national computing centers—it offers immense value for pretraining and post-training DeepSeek-V3-style MoE models. Its deterministic nature also makes it invaluable for on-policy reinforcement learning post-training and rigorous internal ablations, enabling more reliable and reproducible research. This megakernel represents a significant leap in optimizing large-scale MoE model development, pushing the boundaries of what's possible in high-performance AI training.

ai-newsbreakingmarktechpost

Footnotes & Primary References

Related content

AI Leaders Propose SAFE Guidelines for Cybersecurity Transparency

Members of the Open Secure AI Alliance — now more than 120 organizations strong — are developing new guidelines to strengthen agentic AI cybersecurity as the annual Black Hat confe...

Read article

Apple says more ex-employees may have taken confidential data to OpenAI

Apple says its trade secrets investigation into OpenAI has widened. In a new court filing, Apple claims additional former staff may have retained or accessed confidential informati...

Read article

As AI Increases Demands on Memory, Storage Steps Up

Surging AI demands are driving the need for massive datasets and context windows that burst past the confines of system memory.  But rising needs aren’t met by simply adding more s...

Read article