Produck Inference
Models
Four selected public routes for text chat and streaming. Compare the published prices and context displays, then test against your own workload.
ModelInput / 1MOutput / 1MContextSupportedAction
GPT-OSS-120BA text chat route for workloads where input-token spend is a primary consideration.
gpt-oss-120b$0.04
$0.18
131K
Text · Streaming
$0.14
$0.40
262K
Text · Streaming
DeepSeek V4 FlashA text chat route for evaluating very large, approximately stated context.
deepseek-v4-flash$0.09
$0.20
~1M
Text · Streaming
Qwen3 Next 80BA text chat route where generated output length deserves cost attention.
qwen3-next-80b$0.10
$1.10
262K
Text · Streaming
Input and output prices are USD per million tokens. Context displays are public route values; ~1M is approximate.
Start with an API key.
The chat endpoint accepts a public model ID with a Bearer key.