INFRA Signal 170
Meta's MTIA 300 training chip embeds NICs and offloads communication to dedicated engines
Meta's MTIA 300 integrates network interfaces and offloading engines into the chip package to improve training performance for recommendation models.
For engineers training large recommendation models, communication collectives often compete with compute on GPUs, causing over 20% degradation. MTIA 300 offloads these operations to dedicated message engines, keeping compute throughput nearly unaffected. The chip also eliminates the PCIe bottleneck by placing NICs directly in the package.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
MTIA 300 integrates two network chiplets with twelve 800 Gbps RDMA NICs, providing 1.2 TB/s I/O without a PCIe bus.
Sixteen dedicated message engines offload collective communication, delivering over 2.8 TB/s reduction throughput without touching the compute grid.
HCCL compiles collectives into subgraphs for autonomous execution, so the host CPU is uninvolved after work is dispatched.
THE CLUSTER
↗