27B multimodal · OpenAI & Anthropic compatible
Qwen 3.8 27B FP8
Qwen 3.8, 27 billion parameters, served in FP8 with image input, streaming and tool calls. One model resident in memory; a subscription picks the context ceiling it is billed at, and nothing else changes between rungs.
Live capacity
Qwen 3.8 27B FP8
Live numbers are unavailable this second. The commitment behind them does not move: committed shares stay at or under capacity as exact integer arithmetic, checked under a database lock on every checkout — which is why tiers here sell out instead of getting slower.
Model ids
One id per context ceiling. Put the one your subscription is billed at in model:; the suffix is the ceiling.
client = OpenAI(base_url="https://api.partitionlabs.ai/v1")stream = client.chat.completions.create( model="qwen3.8-27b-fp8-4k", # the ceiling you subscribed to messages=[{"role": "user", "content": "..."}], service_tier="flex", # queue order, not speed stream=True,)Context ceilings
Every context ceiling is the same model on the same hardware. What you choose is the longest context your subscription is billed for — nothing else changes between rungs, and a longer ceiling costs more per share unit because it holds more of the card.
Ceilings on this page: 4k · 8k · 16k · 32k · 64k · 128k · 256k.
Up to 4k context
qwen3.8-27b-fp8-4k
Requests up to 4k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$10.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 4k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$40.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 4 concurrent
- Flex
- 4 concurrent
- Standard queue
- 8
- Flex queue
- 16
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 4k tokens
- Batch 24h
- 8 concurrent
- Batch 7d
- 16 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$120.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 12 concurrent
- Flex
- 12 concurrent
- Standard queue
- 24
- Flex queue
- 48
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 4k tokens
- Batch 24h
- 24 concurrent
- Batch 7d
- 48 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$240.00 /mo
The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion
- Active
- 24 concurrent
- Flex
- 24 concurrent
- Standard queue
- 48
- Flex queue
- 96
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 4k tokens
- Batch 24h
- 48 concurrent
- Batch 7d
- 96 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistUp to 8k context
qwen3.8-27b-fp8-8k
Requests up to 8k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$12.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 8k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$48.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 4 concurrent
- Flex
- 4 concurrent
- Standard queue
- 8
- Flex queue
- 16
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 8k tokens
- Batch 24h
- 8 concurrent
- Batch 7d
- 16 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$144.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 12 concurrent
- Flex
- 12 concurrent
- Standard queue
- 24
- Flex queue
- 48
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 8k tokens
- Batch 24h
- 24 concurrent
- Batch 7d
- 48 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$288.00 /mo
The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion
- Active
- 24 concurrent
- Flex
- 24 concurrent
- Standard queue
- 48
- Flex queue
- 96
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 8k tokens
- Batch 24h
- 48 concurrent
- Batch 7d
- 96 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistUp to 16k context
qwen3.8-27b-fp8-16k
Requests up to 16k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$14.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 16k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$56.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 4 concurrent
- Flex
- 4 concurrent
- Standard queue
- 8
- Flex queue
- 16
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 16k tokens
- Batch 24h
- 8 concurrent
- Batch 7d
- 16 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$168.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 12 concurrent
- Flex
- 12 concurrent
- Standard queue
- 24
- Flex queue
- 48
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 16k tokens
- Batch 24h
- 24 concurrent
- Batch 7d
- 48 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$336.00 /mo
The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion
- Active
- 24 concurrent
- Flex
- 24 concurrent
- Standard queue
- 48
- Flex queue
- 96
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 16k tokens
- Batch 24h
- 48 concurrent
- Batch 7d
- 96 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistUp to 32k context
qwen3.8-27b-fp8-32k
Requests up to 32k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$20.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 32k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDuo
Waitlist only$40.00 /mo
2/24 of an always-on agent's GPU load — at least 2h a day of non-stop processing
- Active
- 2 concurrent
- Flex
- 2 concurrent
- Standard queue
- 4
- Flex queue
- 8
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 32k tokens
- Batch 24h
- 4 concurrent
- Batch 7d
- 8 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$80.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 4 concurrent
- Flex
- 4 concurrent
- Standard queue
- 8
- Flex queue
- 16
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 32k tokens
- Batch 24h
- 8 concurrent
- Batch 7d
- 16 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistPro
Waitlist only$160.00 /mo
8/24 of an always-on agent's GPU load — at least 8h a day of non-stop processing
- Active
- 8 concurrent
- Flex
- 8 concurrent
- Standard queue
- 16
- Flex queue
- 32
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 32k tokens
- Batch 24h
- 16 concurrent
- Batch 7d
- 32 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHeavy
Waitlist only$240.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 12 concurrent
- Flex
- 12 concurrent
- Standard queue
- 24
- Flex queue
- 48
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 32k tokens
- Batch 24h
- 24 concurrent
- Batch 7d
- 48 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistMax
Waitlist only$480.00 /mo
The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion
- Active
- 22 concurrent
- Flex
- 24 concurrent
- Standard queue
- 44
- Flex queue
- 96
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 32k tokens
- Batch 24h
- 48 concurrent
- Batch 7d
- 96 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistUp to 64k context
qwen3.8-27b-fp8-64k
Requests up to 64k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$30.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 64k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$120.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 2 concurrent
- Flex
- 2 concurrent
- Standard queue
- 4
- Flex queue
- 8
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 64k tokens
- Batch 24h
- 4 concurrent
- Batch 7d
- 8 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$360.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 6 concurrent
- Flex
- 6 concurrent
- Standard queue
- 12
- Flex queue
- 24
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 64k tokens
- Batch 24h
- 12 concurrent
- Batch 7d
- 24 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$720.00 /mo
The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion
- Active
- 12 concurrent
- Flex
- 12 concurrent
- Standard queue
- 24
- Flex queue
- 48
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 64k tokens
- Batch 24h
- 24 concurrent
- Batch 7d
- 48 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistUp to 128k context
qwen3.8-27b-fp8-128k
Requests up to 128k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$50.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 128k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$200.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 128k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$600.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 3 concurrent
- Flex
- 3 concurrent
- Standard queue
- 6
- Flex queue
- 12
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 128k tokens
- Batch 24h
- 6 concurrent
- Batch 7d
- 12 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$1,200.00 /mo
The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion
- Active
- 6 concurrent
- Flex
- 6 concurrent
- Standard queue
- 12
- Flex queue
- 24
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 128k tokens
- Batch 24h
- 12 concurrent
- Batch 7d
- 24 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistUp to 256k context
qwen3.8-27b-fp8-256k
Requests up to 256k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$90.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 256k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$360.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 256k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$1,080.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 256k tokens
- Batch 24h
- 1 concurrent
- Batch 7d
- 2 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$2,160.00 /mo
The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 256k tokens
- Batch 24h
- 1 concurrent
- Batch 7d
- 2 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.
Beyond one full share. Max is the GPU load of one always-on agent. If you need several, or a pool nothing else runs on, we size it with you rather than from a table.
Talk to usPay as you go
Prepaid credits, priced per million tokens on this model. Same model and the same tokens per second as a subscription — what differs is admission.
Credits run at every context ceiling the model sells. Put the rung you want in model: and the request is admitted up to that ceiling and priced from the model's rate card — the same rule as a subscription.
Pay-as-you-go is best-effort: it fills capacity that reserved tenants are not using, which is most of the time. Under contention a subscriber stays ahead of it — but only inside the share they bought. Past that, their traffic falls below every paying pay-as-you-go request, because the reservation is spent. Nobody is rejected for volume; whoever ranks lower queues.
Priority
queue-jump under contention
- Input
- $0.80
- Cache write
- Free
- Cached read
- $0.10
- Output
- $6.00
per 1M tokens
Standard
the default
- Input
- $0.40
- Cache write
- Free
- Cached read
- $0.05
- Output
- $3.00
per 1M tokens
Flex
same model and speed; last in line
- Input
- $0.30
- Cache write
- Free
- Cached read
- $0.0375
- Output
- $2.25
per 1M tokens
Batch
async; completes within 24h or 7d
- Input
- $0.20
- Cache write
- Free
- Cached read
- $0.025
- Output
- $1.50
per 1M tokens
No context surcharge at launch.