Chat with DeepSeek V4 Flash or Pro
Use V4 Flash for fast, economical conversations, or choose V4-Pro for demanding coding, reasoning, and tool-connected work. Either model can power a private Agent.
Choose DeepSeek V4 Flash or Pro
Chat with either shared model directly, or create a new private Agent. Flash is the faster, lower-cost choice; Pro is for harder coding and reasoning tasks.
Chat now or create an Agent
Start a conversation with the fixed model, or use it to create a private Agent with your own Skill and instructions.
Chat now or create an Agent
Start a conversation with the fixed model, or use it to create a private Agent with your own Skill and instructions.
Release and API status
Current public facts from DeepSeek documentation.
DeepSeek V4 Flash and Pro facts
Compare the shared V4 capabilities, then choose Flash for speed and cost or Pro for harder Agent workflows.
- Model ID
- deepseek-v4-pro
- Context window
- 1M
- Max output
- 384K
- Modalities
- Text
- Reasoning
- Thinking and non-thinking
- Features
- JSON, tools, prefix, FIM
Use deepseek-v4-pro for the Pro variant; deepseek-v4-flash is the lower-cost companion model.
DeepSeek documents 1M context for both V4-Pro and V4-Flash.
DeepSeek lists a maximum output of 384K tokens.
The pricing table documents text chat capabilities, tool calls, JSON output, and FIM beta.
Both V4 models support thinking and non-thinking modes.
DeepSeek lists JSON output, tool calls, chat prefix completion, and FIM completion beta.
Best-fit Agent use cases
Prioritize workloads that benefit from long context and low token cost.
Decision evidence
Provider-published facts that matter for Agent selection.
| Dimension | Official evidence | Decision use |
|---|---|---|
| Agentic capabilities | DeepSeek describes V4-Pro as optimized for agentic coding and integrated with leading coding agents. | Shortlist it for coding and tool-loop workflows, then test on your own acceptance scenarios. |
| Long context | DeepSeek documents 1M context and 384K max output. | Useful when the Agent must hold large repositories or documents in one request. |
| Token price | DeepSeek publishes low cache-hit, cache-miss, and output prices for V4-Pro and V4-Flash. | Measure successful-task cost, especially when repeated long prompts benefit from cache hits. |
DeepSeek release claims are provider-published; validate with your own tasks before production use.
Published API pricing
Rates per 1M tokens from DeepSeek pricing.
| Model | Input | Output |
|---|---|---|
| deepseek-v4-pro | $0.003625 hit / $0.435 miss | $0.87 |
| deepseek-v4-flash | $0.0028 hit / $0.14 miss | $0.28 |
Prices are per 1M tokens. Product prices may vary, and DeepSeek recommends checking the pricing page regularly.
Operational limits
Validate these before switching a production Agent.
Official sources
Primary DeepSeek documentation used for this page.
DeepSeek V4 Flash and Pro FAQ
Answers about DeepSeek V4 online Chat, Flash versus Pro, API pricing, context, and Agent creation.
Start with the right DeepSeek V4 model
Chat with Flash for speed and lower cost, use Pro for harder work, or create a private DeepSeek V4 Agent.

