Produck Inference · Model
GPT-OSS-120B API
A text chat route for workloads where input-token spend is a primary consideration.
Input / 1M$0.04
Output / 1M$0.18
Context131K
Produck model ID
gpt-oss-120bIntegration
API examples
Send a server-side HTTP request to the OpenAI-compatible chat completions endpoint. The examples use the public Produck model ID.
Shell
curl https://api.produckai.com/v1/chat/completions \
-H "Authorization: Bearer $PRODUCK_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model":"gpt-oss-120b",
"messages":[
{"role":"user","content":"Hello"}
]
}'Python · requests
import os
import requests
url = "https://api.produckai.com/v1/chat/completions"
headers = {
"Authorization": f"Bearer {os.environ['PRODUCK_API_KEY']}",
"Content-Type": "application/json",
}
payload = {"model": "gpt-oss-120b", "messages": [{"role": "user", "content": "Hello"}]}
response = requests.post(url, headers=headers, json=payload, timeout=60)
response.raise_for_status()
print(response.json())TypeScript · server-side fetch
// Server-side TypeScript (Node.js)
const response = await fetch("https://api.produckai.com/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.PRODUCK_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "gpt-oss-120b",
messages: [{ role: "user", content: "Hello" }],
}),
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
console.log(await response.json());Workload choice
Why use GPT-OSS-120B?
Use GPT-OSS-120B when you want a clear input-token price for text chat. Evaluate it against your actual request sizes, response lengths, and latency targets before choosing an operating point.
Current API surface
Supported through Produck Inference
- Text chat
- Streaming
Model facts
- Public model ID
gpt-oss-120b- Public context display
- 131K
- API endpoint
https://api.produckai.com/v1/chat/completions
Workload guidance
Match the route to the request.
Good fit for
- Text chat workloads with substantial input volume
- Teams comparing token spend across repeated requests
- Projects that fit within its published context window
Consider another model when
- Your request needs more than its displayed context
- A different input/output price balance better matches your workload
Alternatives
Related models
Inference Engineering