27B multimodal · OpenAI & Anthropic compatible
Qwen 3.8 27B
Qwen 3.8, 27 billion parameters, served with FP8 weights and an FP8 KV cache, with image input, streaming and tool calls. One model resident in memory; a subscription picks the context ceiling it is billed at, and nothing else changes between rungs.
Live capacity
Qwen 3.8 27B
Live numbers are unavailable this second. The commitment behind them does not move: committed shares stay at or under capacity as exact integer arithmetic, checked under a database lock on every checkout — which is why tiers here sell out instead of getting slower.
Model id
qwen3.8-27b
Put this in model:. Your plan's ceiling is the longest request it accepts; credits run up to the widest ceiling we serve. Longer requests are refused (HTTP 400), not truncated and not billed.
Quantization
FP8 weights, FP8 (e4m3) KV cache. Every ceiling, price and capacity figure on this page was measured with this configuration; a change to either would be a new model id, never a silent swap.
client = OpenAI(base_url="https://api.partitionlabs.ai/v1")stream = client.chat.completions.create( model="qwen3.8-27b", # one id; your plan sets the ceiling messages=[{"role": "user", "content": "..."}], service_tier="flex", # queue order, not speed stream=True,)Context ceilings
Every context ceiling is the same model on the same hardware. What you choose is the longest context your subscription reserves — a longer ceiling costs more per share unit because it holds more of the card.
Ceilings on this page: 4k · 8k · 16k · 32k · 64k · 128k · 256k.
Up to 4k context
Requests up to 4k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.
Requests longer than this ceiling are refused on this plan.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$10.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 4k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$40.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 4 concurrent
- Flex
- 4 concurrent
- Standard queue
- 8
- Flex queue
- 16
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 4k tokens
- Batch 24h
- 8 concurrent
- Batch 7d
- 16 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$120.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 12 concurrent
- Flex
- 12 concurrent
- Standard queue
- 24
- Flex queue
- 48
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 4k tokens
- Batch 24h
- 24 concurrent
- Batch 7d
- 48 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$240.00 /mo
The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion
- Active
- 24 concurrent
- Flex
- 24 concurrent
- Standard queue
- 48
- Flex queue
- 96
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 4k tokens
- Batch 24h
- 48 concurrent
- Batch 7d
- 96 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistUp to 8k context
Requests up to 8k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.
Requests longer than this ceiling are refused on this plan.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$12.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 8k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$48.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 4 concurrent
- Flex
- 4 concurrent
- Standard queue
- 8
- Flex queue
- 16
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 8k tokens
- Batch 24h
- 8 concurrent
- Batch 7d
- 16 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$144.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 12 concurrent
- Flex
- 12 concurrent
- Standard queue
- 24
- Flex queue
- 48
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 8k tokens
- Batch 24h
- 24 concurrent
- Batch 7d
- 48 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$288.00 /mo
The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion
- Active
- 24 concurrent
- Flex
- 24 concurrent
- Standard queue
- 48
- Flex queue
- 96
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 8k tokens
- Batch 24h
- 48 concurrent
- Batch 7d
- 96 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistUp to 16k context
Requests up to 16k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.
Requests longer than this ceiling are refused on this plan.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$14.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 16k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$56.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 4 concurrent
- Flex
- 4 concurrent
- Standard queue
- 8
- Flex queue
- 16
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 16k tokens
- Batch 24h
- 8 concurrent
- Batch 7d
- 16 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$168.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 12 concurrent
- Flex
- 12 concurrent
- Standard queue
- 24
- Flex queue
- 48
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 16k tokens
- Batch 24h
- 24 concurrent
- Batch 7d
- 48 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$336.00 /mo
The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion
- Active
- 24 concurrent
- Flex
- 24 concurrent
- Standard queue
- 48
- Flex queue
- 96
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 16k tokens
- Batch 24h
- 48 concurrent
- Batch 7d
- 96 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistUp to 32k context
Requests up to 32k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.
Requests longer than this ceiling are refused on this plan.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$20.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 32k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDuo
Waitlist only$40.00 /mo
2/24 of an always-on agent's GPU load — at least 2h a day of non-stop processing
- Active
- 2 concurrent
- Flex
- 2 concurrent
- Standard queue
- 4
- Flex queue
- 8
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 32k tokens
- Batch 24h
- 4 concurrent
- Batch 7d
- 8 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$80.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 4 concurrent
- Flex
- 4 concurrent
- Standard queue
- 8
- Flex queue
- 16
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 32k tokens
- Batch 24h
- 8 concurrent
- Batch 7d
- 16 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistPro
Waitlist only$160.00 /mo
8/24 of an always-on agent's GPU load — at least 8h a day of non-stop processing
- Active
- 8 concurrent
- Flex
- 8 concurrent
- Standard queue
- 16
- Flex queue
- 32
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 32k tokens
- Batch 24h
- 16 concurrent
- Batch 7d
- 32 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHeavy
Waitlist only$240.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 12 concurrent
- Flex
- 12 concurrent
- Standard queue
- 24
- Flex queue
- 48
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 32k tokens
- Batch 24h
- 24 concurrent
- Batch 7d
- 48 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistMax
Waitlist only$480.00 /mo
The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion
- Active
- 22 concurrent
- Flex
- 24 concurrent
- Standard queue
- 44
- Flex queue
- 96
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 32k tokens
- Batch 24h
- 48 concurrent
- Batch 7d
- 96 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistUp to 64k context
Requests up to 64k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.
Requests longer than this ceiling are refused on this plan.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$30.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 64k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$120.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 2 concurrent
- Flex
- 2 concurrent
- Standard queue
- 4
- Flex queue
- 8
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 64k tokens
- Batch 24h
- 4 concurrent
- Batch 7d
- 8 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$360.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 6 concurrent
- Flex
- 6 concurrent
- Standard queue
- 12
- Flex queue
- 24
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 64k tokens
- Batch 24h
- 12 concurrent
- Batch 7d
- 24 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$720.00 /mo
The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion
- Active
- 12 concurrent
- Flex
- 12 concurrent
- Standard queue
- 24
- Flex queue
- 48
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 64k tokens
- Batch 24h
- 24 concurrent
- Batch 7d
- 48 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistUp to 128k context
Requests up to 128k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.
Requests longer than this ceiling are refused on this plan.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$50.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 128k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$200.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 128k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$600.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 3 concurrent
- Flex
- 3 concurrent
- Standard queue
- 6
- Flex queue
- 12
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 128k tokens
- Batch 24h
- 6 concurrent
- Batch 7d
- 12 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$1,200.00 /mo
The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion
- Active
- 6 concurrent
- Flex
- 6 concurrent
- Standard queue
- 12
- Flex queue
- 24
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 128k tokens
- Batch 24h
- 12 concurrent
- Batch 7d
- 24 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistUp to 256k context
Requests up to 256k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.
Requests longer than this ceiling are refused on this plan.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$90.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 256k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$360.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 256k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$1,080.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 256k tokens
- Batch 24h
- 1 concurrent
- Batch 7d
- 2 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$2,160.00 /mo
The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Reserved context
- 256k tokens
- Batch 24h
- 1 concurrent
- Batch 7d
- 2 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.
Beyond one full share. Max is the GPU load of one always-on agent. If you need several, or a pool nothing else runs on, we size it with you rather than from a table.
Talk to usPay as you go
Prepaid credits, priced per million tokens on this model. Same model and the same tokens per second as a subscription — what differs is admission.
Credits run up to the widest context ceiling the model serves — one model id, no rung to name. A subscription reserves a ceiling: requests longer than your plan's ceiling are refused rather than run on someone else's reservation.
Pay-as-you-go is best-effort: it fills capacity that reserved tenants are not using, which is most of the time. Under contention a subscriber stays ahead of it — but only inside the share they bought. Past that, their traffic falls below every paying pay-as-you-go request, because the reservation is spent. Nobody is rejected for volume; whoever ranks lower queues.
Priority
queue-jump under contention
- Input
- $0.80
- Cache write
- Free
- Cached read
- $0.10
- Output
- $6.00
per 1M tokens
Standard
the default
- Input
- $0.40
- Cache write
- Free
- Cached read
- $0.05
- Output
- $3.00
per 1M tokens
Flex
same model and speed; last in line
- Input
- $0.30
- Cache write
- Free
- Cached read
- $0.0375
- Output
- $2.25
per 1M tokens
Batch
async; completes within 24h or 7d
- Input
- $0.20
- Cache write
- Free
- Cached read
- $0.025
- Output
- $1.50
per 1M tokens
No context surcharge at launch.