Produck Inference · Model

GPT-OSS-120B API

A text chat route for workloads where input-token spend is a primary consideration.

Input / 1M$0.04
Output / 1M$0.18
Context131K
Produck model ID
gpt-oss-120b

Integration

API examples

API Reference

Send a server-side HTTP request to the OpenAI-compatible chat completions endpoint. The examples use the public Produck model ID.

Shell
curl https://api.produckai.com/v1/chat/completions \
  -H "Authorization: Bearer $PRODUCK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model":"gpt-oss-120b",
    "messages":[
      {"role":"user","content":"Hello"}
    ]
  }'
Python · requests
import os
import requests

url = "https://api.produckai.com/v1/chat/completions"
headers = {
    "Authorization": f"Bearer {os.environ['PRODUCK_API_KEY']}",
    "Content-Type": "application/json",
}
payload = {"model": "gpt-oss-120b", "messages": [{"role": "user", "content": "Hello"}]}

response = requests.post(url, headers=headers, json=payload, timeout=60)
response.raise_for_status()
print(response.json())
TypeScript · server-side fetch
// Server-side TypeScript (Node.js)
const response = await fetch("https://api.produckai.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.PRODUCK_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "gpt-oss-120b",
    messages: [{ role: "user", content: "Hello" }],
  }),
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
console.log(await response.json());

Workload choice

Why use GPT-OSS-120B?

Use GPT-OSS-120B when you want a clear input-token price for text chat. Evaluate it against your actual request sizes, response lengths, and latency targets before choosing an operating point.

Pricing

$0.04 input and $0.18 output per million tokens.

Compare all pricing

Current API surface

Supported through Produck Inference

  • Text chat
  • Streaming

Model facts

Public model ID
gpt-oss-120b
Public context display
131K
API endpoint
https://api.produckai.com/v1/chat/completions

Workload guidance

Match the route to the request.

Good fit for

  • Text chat workloads with substantial input volume
  • Teams comparing token spend across repeated requests
  • Projects that fit within its published context window

Consider another model when

  • Your request needs more than its displayed context
  • A different input/output price balance better matches your workload

Inference Engineering

Need a specific p95 TTFT, throughput, or cost target?