Home/Together AI
Together AI cloud platform favicon

Together AI

Together AI is an AI infrastructure platform for serverless and dedicated inference, GPU clusters, fine-tuning, evaluations, model deployment, and open-model development.

Visit website
Together AI Native Cloud platform overview

Product overview

What is Together AI?

Together AI is a full-stack AI cloud for developers and organizations building with open and custom models. Its platform covers model inference, GPU compute, fine-tuning, evaluations, managed storage, and deployment options ranging from serverless APIs to dedicated infrastructure.

Teams can call hosted models without managing servers, reserve throughput for production workloads, launch GPU clusters for larger jobs, or shape a model with their own training data. The service is infrastructure rather than an end-user assistant: customers still design the application, prompts, data pipeline, safeguards, and evaluation process.

How to Use Together AI

  1. Choose a model from the supported model library or identify a custom model to deploy.
  2. Start with serverless inference for prototyping and variable traffic.
  3. Test quality, latency, context handling, and cost on representative requests.
  4. Add evaluations before fine-tuning; only train when prompt, retrieval, or model selection changes are insufficient.
  5. Move predictable, high-volume workloads to provisioned throughput or dedicated deployments when the economics support it.
  6. Monitor token usage, failures, latency, safety behavior, and model-version changes in production.

Do not send secrets, regulated records, or personal data until your organization has reviewed access controls, retention, regional requirements, and the applicable service terms. Treat model output as untrusted data and validate it before executing code, database queries, financial actions, or customer-facing decisions.

Core Features

  • Serverless inference: Access hosted models through APIs without operating model servers.
  • Batch inference: Process asynchronous or large datasets where immediate responses are unnecessary.
  • Provisioned throughput: Reserve production capacity for more predictable performance.
  • Dedicated inference: Run dedicated models or containers for workloads that need isolation or customization.
  • GPU clusters: Use accelerated compute for training, fine-tuning, and demanding AI workloads.
  • Model fine-tuning: Adapt supported models with LoRA, full fine-tuning, or preference-optimization workflows.
  • Evaluations: Compare models and checkpoints using repeatable quality measurements.
  • Developer sandboxes and storage: Provide environments and managed data storage for development workflows.
  • Open-model library: Explore and deploy a range of open-source models for language and generative tasks.

Use Cases

  • Add text generation, extraction, classification, or conversational features to an application.
  • Serve open models behind a production API without building an inference stack.
  • Fine-tune a model for domain terminology, style, structured output, or task behavior.
  • Run offline processing over large document or data collections.
  • Benchmark multiple models for quality, latency, and cost before deployment.
  • Operate dedicated GPU capacity for custom models or sustained workloads.
  • Build internal AI platforms that need a shared inference and evaluation layer.

Pricing

Pricing is usage-based and varies by product, model, token direction, training method, and infrastructure choice. Serverless inference is generally billed per million input and output tokens; fine-tuning is billed by training and evaluation tokens with minimum job charges. Dedicated endpoints, provisioned throughput, and GPU clusters have separate capacity-based pricing. Review the live calculator and model pricing before deployment because rates and available models change.

Frequently Asked Questions

Is Together AI an AI chatbot?

No. It is infrastructure used to build and operate AI applications. A developer or product team supplies the interface, workflow, data, and business logic.

Does it support open-source models?

Yes. Open-model access and deployment are central parts of the platform, alongside support for fine-tuning and custom deployment patterns.

When should a team use dedicated inference?

Dedicated or provisioned capacity is most useful when traffic is sustained, performance must be predictable, or the workload needs a specific model, container, or isolation profile.

Should every project fine-tune a model?

No. First establish an evaluation set and test prompting, retrieval, tool use, and model selection. Fine-tuning is justified when those approaches cannot reliably meet the target behavior.

Back to product directory

Related products

OpenAI Codex is an AI coding product for delegating scoped repository tasks and reviewing the resulting changes.

Claude Code is Anthropic’s AI coding product for iterative repository work through supported development surfaces.

Gemini CLI is Google’s open-source AI agent for terminal workflows and local development tasks.

Newsletter

Keep up with useful AI products

Get a concise selection of new products, practical use cases, and builder updates.