Home/Groq
Groq LPU AI inference platform icon

Groq

Groq is an AI inference platform that combines purpose-built LPU hardware with GroqCloud to serve supported language models at low latency and cost through developer and enterprise infrastructure.

Visit website
GroqCloud serving fast model inference on LPU infrastructure

Product overview

What is Groq?

Groq is an AI inference provider built around its Language Processing Unit, or LPU, a custom processor designed specifically for inference. GroqCloud gives developers and organizations access to supported models running on this LPU-based stack, with an emphasis on predictable speed, low latency, and cost-efficient production workloads.

The platform operates inference infrastructure in data centers across multiple regions. Developers can obtain an API key, select an available model, and integrate generated text or reasoning into an application. Enterprise users can discuss deployment, scale, reliability, and regional requirements with Groq.

How to Use Groq

  1. Create a GroqCloud developer account and obtain an API key.
  2. Review the currently available models and their capabilities.
  3. Choose a model that fits the application's latency, quality, context, and cost needs.
  4. Integrate the inference API into the application or agent workflow.
  5. Measure response quality, throughput, latency, and spend under realistic traffic.
  6. Move to enterprise capacity or deployment support when production requirements grow.

Core Features

  • LPU inference hardware: Uses custom silicon designed around language-model inference.
  • GroqCloud: Provides managed access to supported models on the Groq stack.
  • Low-latency generation: Targets fast responses for interactive and real-time applications.
  • Developer API: Lets software teams add hosted model inference to products and agents.
  • Global infrastructure: Runs LPU-based systems in data centers across regions.
  • Enterprise support: Addresses production scale, reliability, and organizational requirements.

Use Cases

  • Conversational AI: Reduce response delay for chatbots and voice applications.
  • AI agents: Run repeated model calls in tool-using and multi-step workflows.
  • Real-time analysis: Generate decisions or summaries for latency-sensitive operations.
  • Developer products: Add hosted language-model features without operating GPU infrastructure.
  • High-throughput inference: Serve model workloads where speed and unit economics are important.

Pricing

Groq offers a free API key for developers and links to usage-based pricing for supported models. Exact model prices and rate limits change, so developers should verify the live GroqCloud pricing table before estimating production cost.

Frequently Asked Questions

What is an LPU?

It is Groq's custom processor architecture designed specifically for AI inference.

Does Groq train the models it serves?

The homepage focuses on inference infrastructure and access to supported models, not on a single proprietary model family.

Can developers try Groq without an enterprise contract?

Yes. The developer navigation offers a free API key and direct start-building path.

Back to product directory

Related products

OpenAI Codex is an AI coding product for delegating scoped repository tasks and reviewing the resulting changes.

Claude Code is Anthropic’s AI coding product for iterative repository work through supported development surfaces.

Gemini CLI is Google’s open-source AI agent for terminal workflows and local development tasks.

Newsletter

Keep up with useful AI products

Get a concise selection of new products, practical use cases, and builder updates.