Earnings
Home›Earnings›Previews›AMD and Cerebras unveil disaggregated inference archit…
AMD and Cerebras unveil disaggregated inference architecture for faster AI output
The two firms say their setup targets about fivefold more tokens per second per watt than standalone hardware, aiming to improve performance in energy constrained data centers.
AMD and Cerebras Systems have unveiled a disaggregated AI inference architecture that separates prompt processing from token generation, addressing bottlenecks they say emerge with monolithic chip designs during the inference phase. The companies describe their heterogeneous approach as engineered to improve efficiency for ultra low latency workloads tied to enterprise AI demand.
According to the firms, the architecture uses AMD’s Helios rack scale systems for high throughput processing of prompts and context windows, while Cerebras’ Wafer Scale Engine handles token generation token by token. By running these engines as a single disaggregated workflow, the design aims to avoid the memory wait problem that can occur when one monolithic processor tries to do both tasks.
The announcement also frames efficiency as a pricing and capacity lever for cloud operators, noting that data center teams are constrained by power availability and cooling capacity. In that context, the companies say their joint solution is intended to deliver a fivefold increase in tokens per second per watt versus standalone hardware and target 30% more inference tokens per dollar than legacy monolithic racks.
The update comes from a writeup by MarketBeat Ratings tied to the Advancing AI 2026 event. It argues the architecture positions both hardware developers to compete for the premium enterprise inference market by improving output performance within the same energy footprint.