Produck Inference · Model
DeepSeek V4 Flash API
A text chat route for evaluating very large, approximately stated context.
Input / 1M$0.09
Output / 1M$0.20
Context~1M
Produck model ID
deepseek-v4-flashApproximate public context display.
Integration
API examples
Send a server-side HTTP request to the OpenAI-compatible chat completions endpoint. The examples use the public Produck model ID.
Shell
curl https://api.produckai.com/v1/chat/completions \
-H "Authorization: Bearer $PRODUCK_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model":"deepseek-v4-flash",
"messages":[
{"role":"user","content":"Hello"}
]
}'Python · requests
import os
import requests
url = "https://api.produckai.com/v1/chat/completions"
headers = {
"Authorization": f"Bearer {os.environ['PRODUCK_API_KEY']}",
"Content-Type": "application/json",
}
payload = {"model": "deepseek-v4-flash", "messages": [{"role": "user", "content": "Hello"}]}
response = requests.post(url, headers=headers, json=payload, timeout=60)
response.raise_for_status()
print(response.json())TypeScript · server-side fetch
// Server-side TypeScript (Node.js)
const response = await fetch("https://api.produckai.com/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.PRODUCK_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "deepseek-v4-flash",
messages: [{ role: "user", content: "Hello" }],
}),
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
console.log(await response.json());Workload choice
Why use DeepSeek V4 Flash?
DeepSeek V4 Flash is the catalog option with an approximate large-context display. Plan around that approximation, and test actual prompt and output lengths against the production route before committing a workload.
Current API surface
Supported through Produck Inference
- Text chat
- Streaming
Model facts
- Public model ID
deepseek-v4-flash- Public context display
- ~1M
- API endpoint
https://api.produckai.com/v1/chat/completions
Workload guidance
Match the route to the request.
Good fit for
- Workloads evaluating very long text context
- Teams that can plan around an approximate context display
- Projects comparing input and output token economics
Consider another model when
- You need a stated exact context limit before integration
- A smaller displayed context covers your workload
Alternatives
Related models
Inference Engineering