Token pricing
Produck Inference Pricing
Published input and output prices for each public model route. Compare both sides of a request when estimating spend.
ModelInput / 1MOutput / 1MContext
GPT-OSS-120BA text chat route for workloads where input-token spend is a primary consideration.
gpt-oss-120b$0.04
$0.18
131K
$0.14
$0.40
262K
DeepSeek V4 FlashA text chat route for evaluating very large, approximately stated context.
deepseek-v4-flash$0.09
$0.20
~1M
Qwen3 Next 80BA text chat route where generated output length deserves cost attention.
qwen3-next-80b$0.10
$1.10
262K
USD per million tokens. The ~1M context display is approximate.
Token metering
Input and generated output are priced separately. Estimate both prompt and response volume for your workload.
Balance and account limits
Review your available balance and current account limits in the authenticated Produck console before production use.
No GPU provisioning
Call the hosted API with a model ID and your key; the public integration does not require you to provision a GPU.