Produck Inference · Model
Gemma-4-31B API
A text chat route for workloads that need a larger stated context window.
Input / 1M$0.14
Output / 1M$0.40
Context262K
Produck model ID
gemma-4-31bIntegration
API examples
Send a server-side HTTP request to the OpenAI-compatible chat completions endpoint. The examples use the public Produck model ID.
Shell
curl https://api.produckai.com/v1/chat/completions \
-H "Authorization: Bearer $PRODUCK_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model":"gemma-4-31b",
"messages":[
{"role":"user","content":"Hello"}
]
}'Python · requests
import os
import requests
url = "https://api.produckai.com/v1/chat/completions"
headers = {
"Authorization": f"Bearer {os.environ['PRODUCK_API_KEY']}",
"Content-Type": "application/json",
}
payload = {"model": "gemma-4-31b", "messages": [{"role": "user", "content": "Hello"}]}
response = requests.post(url, headers=headers, json=payload, timeout=60)
response.raise_for_status()
print(response.json())TypeScript · server-side fetch
// Server-side TypeScript (Node.js)
const response = await fetch("https://api.produckai.com/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.PRODUCK_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "gemma-4-31b",
messages: [{ role: "user", content: "Hello" }],
}),
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
console.log(await response.json());Workload choice
Why use Gemma-4-31B?
Gemma-4-31B offers a clearly stated context display for text chat and streaming. Compare its input and output rates with your prompt and response volumes, then validate your application workload on the production route.
Current API surface
Supported through Produck Inference
- Text chat
- Streaming
Model facts
- Public model ID
gemma-4-31b- Public context display
- 262K
- API endpoint
https://api.produckai.com/v1/chat/completions
Workload guidance
Match the route to the request.
Good fit for
- Text chat requests that need a larger stated context window
- Applications with long prompts and controlled response lengths
- Teams comparing input and output spend separately
Consider another model when
- A smaller window is sufficient and input price dominates your costs
- Your workload requires more than its displayed context
Alternatives
Related models
Inference Engineering