HomeFull speed ahead. Keep track of advancements in every frontier
LPU and GroqCloudGroq—
LPU and GroqCloud
Photo: Avivweinstein · CC BY-SA 4.0 (opens commons.wikimedia.org)

LPU and GroqCloud

by GroqUS

DeployStage 4 of 5

GroqCloud serves open models online. NVIDIA now also licenses the LPU design.

Updated 24 Dec 2025Checked 25 Sep0 updates this week

Milestones

No announced next step
  1. First GroqChip2020Complete.
  2. GroqCloud public LLM demos go viral for speedFeb 2024Complete.
  3. Saudi Arabia data center in DammamDec 2024Complete.
  4. NVIDIA licensing agreement24 Dec 2025Complete.

Most important updates

  • 24 Dec 2025
  • Sep 2025
  • Jul 2025
  • Feb 2025
  • Aug 2024

Current obstacles

  • Chip count for large modelsEach chip holds little memory, so big models need hundreds of chips, raising cost.

Physics limits

  • Very little memory per chipAbout 230 MB per chip means a 70-billion-parameter model needs hundreds of chips. SRAM density is barely improving on new nodes, so this doesn't fix itself with the next process.
  • Static schedules need predictable workPlanning every step at compile time works only when the computation is known in advance. Variable-length inputs and branching leave scheduled capacity idle.
  • Every hop between chips adds delaySpreading a model over hundreds of chips means each token crosses many links, and each hop adds latency and energy. That caps the speed advantage for the very largest models.

How it works

3 parts
An Intel Xeon die, not Groq: the big regular blocks in the lower half are SRAM, the fast on-chip memory an LPU uses for all its data
An Intel Xeon die, not Groq: the big regular blocks in the lower half are SRAM, the fast on-chip memory an LPU uses for all its dataPhoto: cole8888 · CC BY-SA 2.0 (opens commons.wikimedia.org)
Memory

SRAM only

Each chip holds about 230 MB of on-chip SRAM and no external memory, so reading weights is very fast, but big models span hundreds of chips.

Compiler

Everything planned in advance

The compiler schedules every calculation and data move to the clock cycle, so there are no caches or waiting at run time.

Network

Chips as one assembly line

Chips pass data directly to each other on a fixed schedule, acting like one long pipeline for each token.

Update log

5 updates

Wed 24 Dec 2025

  • Major: BlogPartnership

Sep 2025

  • Minor: PressFunding

Jul 2025

  • Minor: BlogCommercial

Feb 2025

  • Major: BlogCommercial

Aug 2024

  • Minor: PressFunding

About Groq

The team behind LPU and GroqCloud

Groq

LPU inference chips, GroqCloud

Groq, founded in 2016 by ex-Google chip engineer Jonathan Ross, designed the LPU, a chip for running AI models fast. In December 2025 NVIDIA licensed its technology and hired Ross. Groq still runs its GroqCloud service.

  • NVIDIA licensed Groq's inference technology, non-exclusively, for about $20B (Dec 2025).
  • Simon Edwards became CEO after Jonathan Ross joined NVIDIA.
  • Runs GroqCloud data centers, including a large site in Saudi Arabia.
Founded
201610 yrs
Headquarters
United States
Status
Private
Valuation
$6.9BprivateSep 2025
Raised
$1.4B2 rounds
Last round
Venture round · $750MSep 2025$6.9B post
Works in
ComputingAI chipsAI
Coverage
1 program · 5 updateslatest 24 Dec 2025checked 25 Sep
Lead investors
DisruptiveBlackRock
Customers
Saudi Arabia (HUMAIN / Aramco Digital)
Partners
NVIDIA
People
Jonathan RossFounder (joined NVIDIA in 2025)Simon EdwardsCEO
Primer