Token pricing

Produck Inference Pricing

Published input and output prices for each public model route. Compare both sides of a request when estimating spend.

ModelInput / 1MOutput / 1MContext
GPT-OSS-120BA text chat route for workloads where input-token spend is a primary consideration.gpt-oss-120b
$0.04
$0.18
131K
Gemma-4-31BA text chat route for workloads that need a larger stated context window.gemma-4-31b
$0.14
$0.40
262K
DeepSeek V4 FlashA text chat route for evaluating very large, approximately stated context.deepseek-v4-flash
$0.09
$0.20
~1M
Qwen3 Next 80BA text chat route where generated output length deserves cost attention.qwen3-next-80b
$0.10
$1.10
262K

USD per million tokens. The ~1M context display is approximate.

Token metering

Input and generated output are priced separately. Estimate both prompt and response volume for your workload.

Balance and account limits

Review your available balance and current account limits in the authenticated Produck console before production use.

No GPU provisioning

Call the hosted API with a model ID and your key; the public integration does not require you to provision a GPU.

Make your first request.