Per GPU-hour
$5.00/GPU-hr
GraphN / Pricing
Token pricing for hosted models, GPU runtime for custom models, and platform meter rates. Numbers mirror the billing catalog and render fail-closed when a rate is unpublished.
USD per 1M tokens. Use built-ins for token billing, custom deployments for GPU runtime, and imported models through your provider account.
| Model | Input | Output |
|---|---|---|
| qwen3-80b | $0.20/1M | $0.60/1M |
| qwen3-235b | $0.22/1M | $0.88/1M |
| qwen3-coder | $0.90/1M | $0.90/1M |
| nemotron-3-super | $0.90/1M | $0.90/1M |
| gpt-oss-120b | $0.15/1M | $0.60/1M |
| qwen2.5-vl-7b-instruct | $0.05/1M | $0.05/1M |
| qwen3-vl | $0.10/1M | $0.15/1M |
| qwen3.8-27b | $0.45/1M | $3.20/1M |
| qwen3.5-122b-a10b-fp8 | $0.29/1M |
.40/1M
Per GPU-hour
$5.00/GPU-hr
Per GPU-second
$0.00138889/GPU-second
KB embedding tokens
$0.02/1M input tokens
KB rerank tokens
$0.02/1M input tokens
Unstructured RAG ingestion
$4.00/1,000 pages
Markdown conversion
$0.50/1,000 pages
Object storage
$0.069/GB-month
Standing bytes are measured hourly; 730 hours equals one billing month.
Knowledge base vector storage
Qdrant vector and payload bytes are measured hourly.
Function invocation
$0.40/1M calls
Connector invocation
$0.40/1M calls
Function compute
$0.0000166667/GB-second
RAG job trigger
$0/call
Free allowance: $0.00/call
One GPU-second is one second of active runtime on a single GPU for a custom model using your weights on GraphN-managed GPUs. Runtime is billed per GPU, so a deployment on two GPUs accrues two GPU-seconds for every wall-clock second.
No. Built-ins bill by token usage on GraphN. Imported models run on your external provider and are billed by that provider.
In the app, go to your organization billing pricing section for calculator-style estimates and breakdown arithmetic.