Live · 7am IST · DailyFeatured
Reel

The ShiftMaker

AI Intelligence Daily
Featured

NVIDIA GB300 NVL72 Sets World Record for MoE Pre-Training with DeepSeek-V3 671B

The system achieved 1,648 TFLOPs per GPU during training, using 72 GB300 GPUs. This marks a significant milestone in AI hardware and software integration.

Published 21 July 2026 · ID 2026-07-21-nvidia-gb300-nvl72-sets-world-record-for-moe-pre-training-with-deepseek-v3-671b

NVIDIA GB300 NVL72 has set a new world record for pre-training a large-scale AI model using a mixture of experts (MoE) architecture. This achievement highlights the convergence of hardware and software innovations that are reshaping the landscape of AI training. The DeepSeek-V3 671B model, trained on this system, demonstrates the potential of MoE to scale efficiently across thousands of GPUs. This development is a testament to the advancements in AI platform technology, from silicon to networking and software.

The shift toward MoE in large-scale AI training is fundamentally changing the factors that limit model performance. As compute per token decreases, communication efficiency becomes a critical determinant of how well models can scale. NVIDIA GB300 NVL72 has demonstrated that with the right combination of hardware and software, models can achieve levels of performance that were previously unattainable. This system leverages the latest advancements in GPU technology and networking to enable faster and more efficient training processes.

The system achieved a 1,648 TFLOPs per GPU during the pre-training of the DeepSeek-V3 671B model. This level of performance is made possible by the integration of advanced silicon, high-speed networking, and optimized software. The use of 72 GB300 GPUs in this configuration underscores the importance of scalable infrastructure in modern AI training. This achievement sets a new benchmark for what is possible with current GPU technology and highlights the potential for further improvements in the future.

The implications of this achievement are significant for the AI industry. As training efficiency improves, the cost of developing large-scale models may decrease, making AI more accessible to a broader range of organizations. However, the reliance on specialized hardware like the GB300 NVL72 could lead to increased vendor lock-in and raise concerns about governance and data privacy. Market players may need to adapt to these changes by investing in infrastructure that supports efficient and scalable AI training.

While this milestone represents a major step forward in AI training, the technology is still evolving. The continued development of MoE architectures and the integration of new hardware and software innovations will likely shape the future of AI. As the industry moves forward, the focus will remain on optimizing communication efficiency and reducing the barriers to large-scale model training. This progress will have far-reaching consequences for the development and deployment of AI systems.

Sources

Share on X Share on LinkedIn