Produck Inference · Model

Gemma-4-31B API

A text chat route for workloads that need a larger stated context window.

Input / 1M$0.14
Output / 1M$0.40
Context262K
Produck model ID
gemma-4-31b

Integration

API examples

API Reference

Send a server-side HTTP request to the OpenAI-compatible chat completions endpoint. The examples use the public Produck model ID.

Shell
curl https://api.produckai.com/v1/chat/completions \
  -H "Authorization: Bearer $PRODUCK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model":"gemma-4-31b",
    "messages":[
      {"role":"user","content":"Hello"}
    ]
  }'
Python · requests
import os
import requests

url = "https://api.produckai.com/v1/chat/completions"
headers = {
    "Authorization": f"Bearer {os.environ['PRODUCK_API_KEY']}",
    "Content-Type": "application/json",
}
payload = {"model": "gemma-4-31b", "messages": [{"role": "user", "content": "Hello"}]}

response = requests.post(url, headers=headers, json=payload, timeout=60)
response.raise_for_status()
print(response.json())
TypeScript · server-side fetch
// Server-side TypeScript (Node.js)
const response = await fetch("https://api.produckai.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.PRODUCK_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "gemma-4-31b",
    messages: [{ role: "user", content: "Hello" }],
  }),
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
console.log(await response.json());

Workload choice

Why use Gemma-4-31B?

Gemma-4-31B offers a clearly stated context display for text chat and streaming. Compare its input and output rates with your prompt and response volumes, then validate your application workload on the production route.

Pricing

$0.14 input and $0.40 output per million tokens.

Compare all pricing

Current API surface

Supported through Produck Inference

  • Text chat
  • Streaming

Model facts

Public model ID
gemma-4-31b
Public context display
262K
API endpoint
https://api.produckai.com/v1/chat/completions

Workload guidance

Match the route to the request.

Good fit for

  • Text chat requests that need a larger stated context window
  • Applications with long prompts and controlled response lengths
  • Teams comparing input and output spend separately

Consider another model when

  • A smaller window is sufficient and input price dominates your costs
  • Your workload requires more than its displayed context

Inference Engineering

Need a specific p95 TTFT, throughput, or cost target?