Meta Doubles Training Efficiency for Its LLM-Scale Ads Foundation Model GEM
The update brings 20–25% Model FLOPs Utilization (MFU) and scales training FLOPs 4x in 12 months. The model now runs on several thousand of the latest-generation GPUs.
Meta’s Generative Ads Recommendation Model (GEM) has achieved a significant leap in training efficiency, doubling end-to-end (E2E) training performance to 20–25% Model FLOPs Utilization (MFU). This advancement is part of a broader effort to scale GEM’s capabilities, which underpin ads recommendations on platforms like Instagram and Facebook. The model now operates on several thousand of the latest-generation GPUs, reflecting Meta’s commitment to leveraging cutting-edge hardware for AI development.
The improvements were made possible through a combination of advanced techniques, including Fully Sharded 2D Model Parallelism (FSDP), Load Balancing, and AutoAC. These innovations helped optimize resource allocation and reduce training bottlenecks. Meta also integrated Generalized Dot-Product Attention and Flash Attention to enhance computational efficiency, ensuring that GEM can process vast amounts of data more effectively.
By 2025, Meta aims to scale training FLOPs 4x in 12 months, a target that aligns with the company’s broader vision of advancing AI capabilities at LLM scale. The use of technologies like NCCLX and MXFP8 has played a pivotal role in achieving these performance gains. This progress is expected to significantly impact the speed and accuracy of ad recommendations, ultimately enhancing user engagement and advertiser outcomes.
The increased efficiency and scalability of GEM could lead to reduced training costs and faster deployment of AI models. However, it also raises questions about vendor lock-in, as reliance on specific hardware and software frameworks may limit flexibility. Additionally, governance and regulatory considerations will become more critical as the model’s influence expands across Meta’s platforms and beyond.
While the current developments mark a major milestone, the model is still evolving. Meta continues to refine GEM’s architecture and training methodologies, aiming for further improvements in the coming years. The focus remains on balancing performance gains with sustainable, long-term AI development strategies.