ScitiX Model Inference Pricing
Transparent per-token pricing for curated frontier models on ScitiX bare-metal B200/H100 clusters. All prices in USD. OpenAI-compatible API. 21 models available.
Token pricing — USD per 1M tokens
| Model | Provider | Type |
Input /1M | Output /1M | Cache hit /1M | Context |
| deepseek-ai/deepseek-v4-flash-0731 |
DeepSeek |
TextGeneration |
$0.14 |
$0.28 |
$0.028 |
1M |
| tencent/hy3 |
Tencent |
TextGeneration |
$0.14 |
$0.58 |
$0.035 |
256K |
| glm-5.2 |
Z.ai |
TextGeneration |
$1.40 |
$4.40 |
$0.26 |
1M |
| Qwen/Qwen3-Embedding-8B |
Qwen |
Embedding |
$0.040 |
— |
— |
32K |
| kimi-k2.6 |
Moonshot AI |
TextGeneration |
$0.95 |
$4.00 |
$0.16 |
262K |
| DeepSeek-V4-Flash |
Deepseek |
TextGeneration |
$0.14 |
$0.28 |
$0.028 |
1M |
| DeepSeek-V4-Pro |
Deepseek |
TextGeneration |
$1.74 |
$3.48 |
$0.14 |
1M |
| Qwen/Qwen3-Embedding-4B |
Qwen |
Embedding |
$0.020 |
— |
— |
32K |
| IEITYuan/Yuan-embedding-2.0-en |
IEITYuan |
Embedding |
$0.100 |
— |
— |
32K |
| Qwen/Qwen3-Embedding-0.6B |
Qwen |
Embedding |
$0.010 |
— |
— |
32K |
| Qwen/Qwen3.6-27B |
Qwen |
TextGeneration |
$0.060 |
$0.40 |
$0.030 |
262K |
| MiniMaxAI/MiniMax-M2.7 |
Minimax |
TextGeneration |
$0.30 |
$1.20 |
$0.060 |
196K |
| google/gemma-4-31B-it |
Google |
TextGeneration |
$0.100 |
$0.40 |
$0.016 |
131K |
| Qwen/Qwen3.5-397B-A17B |
Qwen |
TextGeneration |
$0.48 |
$2.88 |
$0.24 |
262K |
| MiniMaxAI/MiniMax-M2.5 |
Minimax |
TextGeneration |
$0.30 |
$1.20 |
$0.030 |
196K |
| google/gemma-3-27b-it |
Google |
TextGeneration |
$0.080 |
$0.16 |
$0.032 |
131K |
| Qwen/Qwen3-32B |
Qwen |
TextGeneration |
$0.040 |
$0.32 |
$0.032 |
40K |
| openai/gpt-oss-120b |
Openai |
TextGeneration |
$0.15 |
$0.40 |
$0.080 |
131K |
| Qwen/Qwen2.5-VL-72B-Instruct |
Qwen |
Vision |
$0.46 |
$0.46 |
$0.32 |
32K |
Speech pricing — USD per minute of audio
| Model | Provider | Type | Price /min audio |
| bosonai/asr |
Boson AI |
AutomaticSpeechRecognition |
$0.006 |
| bosonai/tts |
Boson AI |
TextToSpeech |
$0.050 |
Batch API is billed at 50% of the listed price on eligible models. Cache-hit price applies to cached input tokens.