AI Policy Wiki
Dashboard

Specialty Inference Hardware — Cerebras, Groq, SambaNova

high confidence · updated 2026-06-26

Non-GPU AI accelerator startups: Cerebras wafer-scale CS-3, Groq LPU, SambaNova RDU. Differentiated from Nvidia's training-dominated stack by focusing on inference throughput and latency.

Cerebras, Groq, and SambaNova are three US fabless startups that build AI accelerators architecturally distinct from Nvidia's GPUs, optimized for inference (serving models) rather than training. Each departs from the paradigm of many discrete chips connected over PCIe, NVLink, or InfiniBand, instead making a different trade-off between silicon area, memory placement, and interconnect. Combined, they account for a small share of AI compute spend, and each markets non-GPU architectures as serving frontier-scale models with better latency or throughput per dollar on inference workloads.

Cerebras — Wafer-Scale CS-3

FieldValue
Founded2016
HQSunnyvale, CA
CEOAndrew Feldman
FlagshipCS-3 (Wafer-Scale Engine 3, WSE-3)

The WSE-3 is fabricated as a single silicon die covering an entire 300mm wafer, with approximately 900,000 cores and roughly 44GB of on-chip SRAM on one piece of silicon, an approach that eliminates most chip-to-chip interconnect bottlenecks. Cerebras's cloud service, "Cerebras Inference," reported record inference tokens-per-second on Llama and other open-weight models through 2024–2025, marketed as the fastest inference in the industry for compatible model sizes. Its customer base includes G42 (a UAE-based anchor investor and customer), Mayo Clinic, national labs, and some sovereign AI deployments. Cerebras filed for an IPO in 2024.

Groq — Language Processing Unit (LPU)

FieldValue
Founded2016
HQMountain View, CA
CEODoug Wightman (2026; co-founder Jonathan Ross departed for Nvidia)
FlagshipLPU — deterministic, compiler-scheduled tensor processor

Groq's LPU replaces the dynamic scheduling and cache hierarchy of GPUs with a fully compiler-scheduled, deterministic dataflow architecture in which every cycle's behavior is known at compile time. Groq Cloud demonstrated first-token latency below 0.5 seconds on 70B-class open-weight models, roughly an order of magnitude lower than contemporary GPU inference for comparable models. In 2024 Groq announced a buildout with Saudi Aramco Digital / HUMAIN for regional AI infrastructure, alongside similar deals with other sovereign customers. The LPU uses on-die SRAM rather than HBM stacks, sidestepping the HBM supply crunch documented in (Source: Raw Sources/SemiAnalysis - CoWoS and HBM Supply Chain.md).

Groq's leadership and strategy shifted in 2026. In late 2025, Nvidia signed a non-exclusive license for Groq's LPU technology and hired away co-founder Jonathan Ross and president Sunny Madra, in a transaction reported at roughly $20 billion and characterized as a license-plus-talent "not-acqui-hire" rather than an acquisition. On June 22, 2026, now led by CEO Doug Wightman, Groq confirmed a $650 million funding round led by Disruptive and Infinitum and said it was pivoting toward its neocloud inference-hosting business — 13 data centers serving more than five million developers — and had hired Alan Rice, formerly of xAI, as chief operating officer and Sinclair Schuller as chief technology officer (Source: techcrunch.com).

SambaNova Systems — Reconfigurable Dataflow Unit (RDU)

FieldValue
Founded2017
HQPalo Alto, CA
CEORodrigo Liang
FlagshipSN40L RDU

SambaNova's RDU is a coarse-grained reconfigurable array, in which compute units and memory are arranged in a spatial dataflow graph reconfigured per model. The SN40L pairs RDU compute with HBM and large DDR pools, allowing single-node deployment of trillion-parameter models that would otherwise require multi-GPU systems. SambaNova targets on-premises enterprise and government deployments, particularly where data cannot leave a facility, rather than competing head-on in hyperscaler cloud.

Role in the AI compute supply chain

All three companies target inference specifically — throughput, latency, or single-node model-size advantages — rather than attempting to match Nvidia across training. They still depend on TSMC for fabrication: Cerebras uses the full wafer, while Groq and SambaNova use more conventional reticle-limited dies. Cerebras and SambaNova also depend on HBM from SK Hynix, Samsung, and Micron. None has a software stack approaching CUDA's ecosystem depth; each relies on importing models from PyTorch and Hugging Face via custom compilers.

Geopolitical relevance

All three are subject to US export controls on advanced AI compute, and exports to China are effectively barred at flagship performance levels. Their customer bases skew toward Gulf-state and other non-US sovereign AI buildouts, including G42 and HUMAIN, making them actors in AI Sovereignty debates. For US AI resilience, viable alternatives in Cerebras, Groq, and SambaNova represent a supply chain less dependent on Nvidia alone.

Market position

Combined, the three companies account for a low single-digit share of AI compute revenue; Nvidia and AMD together remain more than 95% of the merchant accelerator market. None is yet profitable at scale, and all depend on sovereign- or hyperscaler-anchor deals. Each retains technology differentiation distinct from the GPU paradigm.

Relationships