Back to Newsroom

Tencent Open-Sources AngelSpec: A Unified Training Framework for MTP and Block-Parallel Speculative Decoding on Hy3 Models

By Modelverse Editorial·July 30, 2026·2 min read
Tencent Open-Sources AngelSpec: A Unified Training Framework for MTP and Block-Parallel Speculative Decoding on Hy3 Models

Tencent has unveiled AngelSpec, an innovative open-source, PyTorch-native framework designed to revolutionize the training of speculative decoding draft models. Moving beyond the conventional approach of seeking a single, averaged drafter, AngelSpec prioritizes real-world workload heterogeneity. It offers a unified framework supporting both autoregressive multi-token prediction (MTP) and the block-parallel DFlash family, allowing developers to tailor draft models precisely to their specific application needs.

Speculative decoding accelerates inference by having a lightweight drafter propose future tokens, which the larger target model then verifies in a single pass. AngelSpec recognizes that the efficiency of this process—how many tokens are accepted and how long the round takes—varies significantly across domains. For high-entropy tasks like open-ended conversation, where many continuations are valid, MTP's shorter candidate sequences are more effective. Conversely, for low-entropy tasks such as code generation or mathematical reasoning, where longer predictable spans exist, block-parallel drafting excels. To address this, AngelSpec provides two specialized drafters: an MTP model trained on diverse conversational data and a block-diffusion model strengthened with code and mathematics samples.

Crucially, AngelSpec tackles the common train-inference mismatch seen in prior models, where errors accumulate during recurrent unrolling. It implements a shared-parameter, multi-depth scheme where a single MTP block is autoregressively unrolled for multiple steps during training, with predictions from one depth feeding the next. This "Training-Time Test" principle, inspired by EAGLE-3, ensures that the model is robustly prepared for long self-generated chains. By sharing parameters while maintaining depth-specific supervision, AngelSpec significantly improves acceptance rates at deeper draft positions, offering researchers and developers a powerful tool to train more efficient, reliable, and specialized speculative decoding models for a wide array of AI applications.

ai-newsbreakingmarktechpost

Footnotes & Primary References

Related content

A fundamental flaw leaves LLMs strikingly vulnerable to attack

It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the In...

Read article

Best in Class: Stream PC Games and Study on the Same Laptop With GeForce NOW

Back to school means balancing assignments, deadlines and downtime. GeForce NOW makes it easy to have it all. With cloud gaming, everyday laptops used for class can also become GeF...

Read article

Dili raises $21.7M to bring AI compliance to the infrastructure boom

The Series A was led by Khosla Ventures, with participation from Allianz, Rebel Fund, Brick and Mortar Ventures’ Darren Bechtel, and Y Combinator’s Garry Tan.

Read article
© 2026 Modelverse®. All rights reserved.Modelverse Newsroom