Product overview
What is Groq?
Groq is an AI inference provider built around its Language Processing Unit, or LPU, a custom processor designed specifically for inference. GroqCloud gives developers and organizations access to supported models running on this LPU-based stack, with an emphasis on predictable speed, low latency, and cost-efficient production workloads.
The platform operates inference infrastructure in data centers across multiple regions. Developers can obtain an API key, select an available model, and integrate generated text or reasoning into an application. Enterprise users can discuss deployment, scale, reliability, and regional requirements with Groq.
How to Use Groq
- Create a GroqCloud developer account and obtain an API key.
- Review the currently available models and their capabilities.
- Choose a model that fits the application's latency, quality, context, and cost needs.
- Integrate the inference API into the application or agent workflow.
- Measure response quality, throughput, latency, and spend under realistic traffic.
- Move to enterprise capacity or deployment support when production requirements grow.
Core Features
- LPU inference hardware: Uses custom silicon designed around language-model inference.
- GroqCloud: Provides managed access to supported models on the Groq stack.
- Low-latency generation: Targets fast responses for interactive and real-time applications.
- Developer API: Lets software teams add hosted model inference to products and agents.
- Global infrastructure: Runs LPU-based systems in data centers across regions.
- Enterprise support: Addresses production scale, reliability, and organizational requirements.
Use Cases
- Conversational AI: Reduce response delay for chatbots and voice applications.
- AI agents: Run repeated model calls in tool-using and multi-step workflows.
- Real-time analysis: Generate decisions or summaries for latency-sensitive operations.
- Developer products: Add hosted language-model features without operating GPU infrastructure.
- High-throughput inference: Serve model workloads where speed and unit economics are important.
Pricing
Groq offers a free API key for developers and links to usage-based pricing for supported models. Exact model prices and rate limits change, so developers should verify the live GroqCloud pricing table before estimating production cost.
Frequently Asked Questions
What is an LPU?
It is Groq's custom processor architecture designed specifically for AI inference.
Does Groq train the models it serves?
The homepage focuses on inference infrastructure and access to supported models, not on a single proprietary model family.
Can developers try Groq without an enterprise contract?
Yes. The developer navigation offers a free API key and direct start-building path.


