Docs · Streaming
Streaming chat completions
Request streamed text output through the same OpenAI-compatible chat completions endpoint. The examples use ordinary server-side HTTP clients.
Set stream to true
Use the same Bearer key, model ID, and messages array as a non-streaming request. Add "stream": true to the JSON body and read the response incrementally.
HTTP examples
The examples print the streamed response as it arrives; adapt parsing to your server-side application.
Shell
curl -N https://api.produckai.com/v1/chat/completions \
-H "Authorization: Bearer $PRODUCK_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model":"gpt-oss-120b",
"messages":[
{"role":"user","content":"Hello"}
],
"stream":true
}'Python · requests
import os
import requests
url = "https://api.produckai.com/v1/chat/completions"
headers = {
"Authorization": f"Bearer {os.environ['PRODUCK_API_KEY']}",
"Content-Type": "application/json",
}
payload = {"model": "gpt-oss-120b", "messages": [{"role": "user", "content": "Hello"}], "stream": True}
with requests.post(url, headers=headers, json=payload, stream=True, timeout=60) as response:
response.raise_for_status()
for line in response.iter_lines(decode_unicode=True):
if line:
print(line)TypeScript · server-side fetch
// Server-side TypeScript (Node.js)
const response = await fetch("https://api.produckai.com/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.PRODUCK_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "gpt-oss-120b",
messages: [{ role: "user", content: "Hello" }],
stream: true,
}),
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
if (!response.body) throw new Error("Missing response body");
const reader = response.body.getReader();
const decoder = new TextDecoder();
while (true) {
const { value, done } = await reader.read();
if (done) break;
process.stdout.write(decoder.decode(value, { stream: true }));
}Next steps
Compare the available routes and review the supported request fields. Measure TTFT and output cadence using your actual workload before setting production SLOs.