Docs · Streaming

Streaming chat completions

Request streamed text output through the same OpenAI-compatible chat completions endpoint. The examples use ordinary server-side HTTP clients.

Set stream to true

Use the same Bearer key, model ID, and messages array as a non-streaming request. Add "stream": true to the JSON body and read the response incrementally.

HTTP examples

The examples print the streamed response as it arrives; adapt parsing to your server-side application.

Shell
curl -N https://api.produckai.com/v1/chat/completions \
  -H "Authorization: Bearer $PRODUCK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model":"gpt-oss-120b",
    "messages":[
      {"role":"user","content":"Hello"}
    ],
    "stream":true
  }'
Python · requests
import os
import requests

url = "https://api.produckai.com/v1/chat/completions"
headers = {
    "Authorization": f"Bearer {os.environ['PRODUCK_API_KEY']}",
    "Content-Type": "application/json",
}
payload = {"model": "gpt-oss-120b", "messages": [{"role": "user", "content": "Hello"}], "stream": True}

with requests.post(url, headers=headers, json=payload, stream=True, timeout=60) as response:
    response.raise_for_status()
    for line in response.iter_lines(decode_unicode=True):
        if line:
            print(line)
TypeScript · server-side fetch
// Server-side TypeScript (Node.js)
const response = await fetch("https://api.produckai.com/v1/chat/completions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.PRODUCK_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "gpt-oss-120b",
    messages: [{ role: "user", content: "Hello" }],
    stream: true,
  }),
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
if (!response.body) throw new Error("Missing response body");
const reader = response.body.getReader();
const decoder = new TextDecoder();
while (true) {
  const { value, done } = await reader.read();
  if (done) break;
  process.stdout.write(decoder.decode(value, { stream: true }));
}

Next steps

Compare the available routes and review the supported request fields. Measure TTFT and output cadence using your actual workload before setting production SLOs.