Wednesday, May 20, 2026
S&P 500 · NDX · NVDA · Gold · BTC
  Tool · Compute & Inference All tools →
Groq

Groq.

LPU inference hardware — the fastest tokens-per-second in the industry.

Category
Compute & Inference
Vendor
Groq
Domain
groq.com
Pricing
Pay-per-token inference · GroqCloud · On-prem LPU racks
Region
Mountain View, US
Models
Llama, Mistral, Qwen, Whisper — selected open models served on LPU

What it is.

Mountain View hardware company founded by the founder, the engineer who built Google's original TPU. The Language Processing Unit is a custom inference chip that serves selected open models at 500 to 1,200 tokens per second per user — roughly an order of magnitude above NVIDIA H100 baselines. GroqCloud is the hosted API; on-prem LPU racks ship to financial-services and defence buyers.

Where it fits.

Anywhere latency is the product. Voice agent backbones that need sub-300ms first-token responses. Reasoning chains where the user is waiting on every step. Buyers prepared to constrain model choice to the supported set in exchange for speed nothing else can match. On-prem LPU has quiet adoption in US government and tier-one banks.

Use cases
  • Latency-critical chatbots and copilots
  • Real-time voice agent backbones
  • Speculative decoding workloads
  • On-prem inference for regulated buyers
Strengths
  • Tokens-per-second leadership unmatched by GPU competitors
  • Predictable latency under load — no thermal degradation
  • On-prem LPU available for sovereign and regulated deployment
Considerations
  • Model catalog narrower than GPU-based competitors
  • Custom hardware lock-in limits future model portability

Frequently asked.

How fast is Groq compared to GPU inference?

Five to ten times faster on tokens-per-second for supported models. Llama 4 Scout runs over 1,000 tokens per second on LPU versus 150 on H100 reference setups.

Which models run on Groq?

Selected Llama, Mistral, Qwen, and Whisper variants. Adding a new model requires Groq engineering to compile it for the LPU, which is why the catalog is narrower.

Can I buy LPU hardware?

Yes on Enterprise. Groq sells LPU racks for on-prem deployment to financial-services, defence, and government buyers. Pricing is negotiated per cluster.