AMD has partnered with Cerebras Systems to develop a disaggregated compute platform [1], combining AMD's Instinct GPUs with Cerebras' SRAM-powered AI accelerators. This collaboration aims to deliver ultra-low-latency inference for AI workloads, directly challenging Nvidia's Groq LPUs [2]. Cerebras' wafer-scale engines store entire models in on-chip SRAM, providing ultra-high bandwidth and generating tokens without shuttling weights from external HBM, which slows GPUs [3]. The partnership between AMD and Cerebras is expected to boost the number of tokens per second generated per watt of electricity consumed by as much as 5x [4].

This move positions AMD and Cerebras as strong competitors against Nvidia in the AI chip market, with Cerebras' technology offering a unique advantage in terms of price-performance, delivering up to a 6x price-performance advantage over Groq [5]. The combined offering will be available in Cerebras Cloud later this year, marking a significant step in the evolution of AI infrastructure [6].

Sources

  1. https://www.theregister.com/2025/12/31/why_did_nvidia_really_drop_20b_on_groq/
  2. https://www.cerebras.ai/press-release/amd-and-cerebras-announce-industry-leading-ultra-low-latency-and-high-throughput-ai-inference
  3. https://www.sdxcentral.com/analysis/cerebras-spins-nvidias-groq-tieup-as-proof-its-waferscale-bet-was-right/
  4. https://www.cerebras.ai/blog/cerebras-cs-3-vs-groq-lpu
  5. https://intuitionlabs.ai/articles/cerebras-vs-sambanova-vs-groq-ai-chips
  6. https://www.theregister.com/2025/10/13/cerebras_aims_deploy_ai_infrastructure_massive_stargate_uae_data_centre_hub/