AMD and Cerebras Announce Industry-Leading Ultra-Low-Latency and High Throughput AI Inference Solution
News Highlights AMD and Cerebras are collaborating to advance a workload-optimized approach to ultra-low-latency AI inference infrastructure. AMD Helios™ and the Cerebras Wafer-Scale Engine will operate as a single disaggregated inference workflow, combining ultra-high-throughput from AMD Instinct™ GPUs, with ultra-fast token generation of Cerebras Wafer-Scale Engine. Cerebras plans to deploy AMD Helios in its data centers, with the […]
\