Meta unveils MTIA 400 AI chip for LLM training and ad serving
Meta introduced its MTIA 400 accelerator, a custom chip built for large-language-model training while also handling ad-recommender inference, delivering up to 12 petaFLOPS of compute.
Meta has revealed its fourth-generation MTIA 400 accelerator, a heterogeneous multi-die chip built on a 3 nm process that targets both large-language-model training and ad-recommender inference. Two compute dies, each with a 6 × 8 grid of processing elements, together deliver 12 petaFLOPS of MXFP4 performance at 1.7 GHz, while eight 36 GB HBM3e stacks provide 288 GB of memory and about 9.2 TB/s bandwidth. Compared with Nvidia's Blackwell GPUs, the MTIA 400 is roughly 20 percent faster for high-precision training at similar power, though it trails newer Nvidia and AMD GPUs on inference speed.
The chip is packaged with I/O dies offering 1.2 TB/s RDMA bandwidth and is assembled into racks that hold 72 accelerators across 18 compute blades. Meta envisions scaling to multi-thousand-accelerator clusters, and it is already developing the MTIA 450, which will double memory bandwidth with HBM4, and the MTIA 500, slated for 2027, promising further bandwidth and compute gains.
Why it matters
It shows Meta is building its own AI hardware to cut costs and boost performance for both model training and ad serving.
In this story
