AMD's acquisition of Taalas brings model-specific silicon that delivers 16,000 tokens per second per user — multiples beyond current GPU inference.
AMD's acquisition of Taalas brings model-specific silicon that delivers 16,000 tokens per second per user — multiples beyond current GPU inference.

AMD acquired Taalas, whose hardwired AI chips deliver 16,000 tokens per second per user, giving the chipmaker a path to challenge Nvidia in the fast-growing inference market.
"We founded Taalas to rethink AI inference from the ground up by building the hardware around the model," Ljubisa Bajic, CEO and co-founder of Taalas, said. "Joining AMD will give us the scale, engineering resources, and global reach to accelerate our innovation."
Taalas, founded in 2023 by former Tenstorrent leaders including ex-CEO Bajic, emerged from stealth in February with a demo chip achieving more than 16,000 tokens per second per user on Llama3.1-8B — multiples beyond what current-generation GPUs can deliver. The startup's HC1 chip hardwires a model's dataflow into silicon and burns in the weights, borrowing from structured ASIC approaches of the early 2000s. The trade-off: the chip only runs Llama3.1-8B, and switching models requires a new tape-out — typically two masks representing weights and dataflow. The company raised $50 million from Quiet Capital and Pierre Lamond, followed by $169 million from investors including Fidelity earlier this year.
AMD plans to integrate Taalas technology into system-level solutions alongside its Instinct GPUs, mirroring Nvidia's recent all-but-acquisition of Groq and AMD's existing partnership with Cerebras on disaggregated inference. The deal, subject to closing conditions and regulatory approvals, gives AMD a differentiated path into the inference market where efficiency matters more than flexibility.
The Structured-ASIC Bet Behind 16,000 Tokens
Taalas's approach strips away programmability in exchange for speed. The HC1 chip is entirely SRAM-based, with model weights and dataflow burned into the silicon. For larger models, the constraint becomes physical: DeepSeek-671B would require roughly 30 separate tape-outs. The company's tool flow, which converts trained models into mask designs in about two months, is the core of its advantage — a stark contrast to the one-to-two-year development cycles typical of custom ASICs.
The limit for a single Taalas chip appears to be around eight billion parameters, depending on how aggressively the model has been quantized. That makes the technology suitable for small-model inference workloads — edge applications, physical AI, and latency-sensitive deployments where power constraints and cost sensitivity outweigh the need for frequent model updates.
AMD's Inference Stack vs. Nvidia-Groq
AMD's acquisition follows a pattern of disaggregated inference architectures taking shape across the industry. Nvidia all but acquired Groq late last year, planning to use its chips for the decode stage of inference alongside GPUs. AMD recently announced a similar arrangement with Cerebras, where its GPUs handle prefill and Cerebras's wafer-scale engine accelerates decode. A Taalas SRAM-based chip could potentially fill that decode role in-house, giving AMD a fully controlled system analogous to the Nvidia-Groq pairing.
The deal also deepens AMD's Canadian footprint. The company has maintained operations in Toronto since acquiring ATI Technologies in 2006, and this marks its second acquisition of a Canadian AI chip firm in just over a year, following the Untether AI team deal in 2025. AMD is also an investor in Toronto-based Cohere and Xanadu.
AMD shares, trading at $489.28 with returns above 100 percent over the past year, reflect the market's optimism about its AI data center trajectory. Q2 2026 revenue reached $11.5 billion, with the Core Scientific partnership pointing to large future deployments. The Taalas acquisition adds inference-specific silicon to a portfolio that already spans Ryzen AI in consumer PCs, EPYC processors in data centers, and Instinct GPUs for training workloads.
The key question for investors is whether AMD can convert Taalas technology into concrete system wins — inference-optimized chips shipping alongside Instinct GPUs in enterprise or cloud deployments. Evidence would come in the form of specific customer announcements or quantified contributions from inference solutions in future data center updates. Nvidia, with its CUDA software platform and aggressive competitive responses, remains the benchmark against which any AMD inference strategy will be judged.
This article is for informational purposes only and does not constitute investment advice.