INFRA Signal 532 2 feeds carried it
Intel details Crescent Island AI accelerator with 16-deep XMX engines and expanded caches for inference-first workloads
Intel disclosed architectural details of its Crescent Island AI accelerator at Hot Chips 2026, positioning the 350W air-cooled PCIe card as an inference-first alternative to higher-power competitors from Nvidia and AMD.
Crescent Island deploys in traditional air-cooled servers with LPDDR5X memory rather than requiring liquid cooling and HBM4, lowering the infrastructure bar for AI inference. Its 16-deep XMX systolic engines and doubled register file space are specifically tuned for mixture-of-experts models paired with speculative decoding, a workload class Intel sees growing.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Crescent Island uses four Xe3P slices with 32 Xe Cores total, each containing 1MB of general register file space (double Battlemage) and 512KB of L1 cache or shared local memory.
The XMX matrix accelerators use a 16-deep systolic design, up from 4-deep on Xe2 and Xe3, allowing larger matrix chunks per operation during general matrix-multiply workloads.
The chip supports data types from MXFP4 to full-rate FP64 via 64 FP64 FMA units per Xe Core, and includes sigmoid and tanh transcendental functions important for softmax during inference.
THE CLUSTER
↗