AI factories operate continuously and their economics hinge on measurable output such as tokens processed per second, energy efficiency, cost per token, system utilization and uptime. Achieving these metrics requires treating the AI infrastructure as an integrated factory rather than a set of discrete accelerators. Hyperscalers and AI‑native firms that design custom XPUs must therefore consider the entire platform—scale‑up and scale‑out networking, rack‑level architecture, factory‑grade software, and a reliable supplier ecosystem. At the scale of an AI factory this end‑to‑end effort becomes both complex and expensive, slowing the introduction of new XPUs to market.
NVLink Fusion addresses this bottleneck by linking custom XPUs to NVIDIA’s established AI infrastructure. The connection boosts performance, shortens time‑to‑market and reduces risk for semi‑custom AI factories. For workloads that demand trillion‑parameter models, mixture‑of‑experts layouts or agentic AI, a scale‑up fabric that cannot keep pace drives down utilization and raises the cost per token. Consequently, a scale‑up networking solution must satisfy three core dimensions: bandwidth, latency, and reliability.
- NVLink Fusion provides high‑bandwidth, low‑latency interconnect between XPUs and NVIDIA’s AI stack.
- Enables support for trillion‑parameter models, MoE architectures and agentic AI workloads.
- Improves system utilization and lowers cost per token by keeping the scale‑up fabric matched to compute demand.
- Accelerates XPU time‑to‑market while leveraging proven NVIDIA infrastructure and software.
- Reduces integration risk for semi‑custom AI factories by reusing mature networking and rack‑scale design.
Why this matters
The source notes that building a full AI factory is a “fundamental obstacle” to getting XPUs to market quickly because of the complexity and cost of designing scale‑up networking, rack architecture and software. NVLink Fusion directly tackles that obstacle by offering a ready‑made, high‑performance interconnect that lets XPU designers concentrate on the accelerator itself while reusing NVIDIA’s validated infrastructure. This separation of concerns can shorten development cycles and improve the economic viability of semi‑custom AI factories, as evidenced by the promised gains in utilization and cost per token.
Share this article
Found this insightful? Share it with your community on Reddit, X, or copy the link.
