27B multimodal · OpenAI & Anthropic compatible

Qwen 3.8 27B

Qwen 3.8, 27 billion parameters, served with FP8 weights and an FP8 KV cache, with image input, streaming and tool calls. One model resident in memory; a subscription picks the context ceiling it is billed at, and nothing else changes between rungs.

Live capacity

Qwen 3.8 27B

Live numbers are unavailable this second. The commitment behind them does not move: committed shares stay at or under capacity as exact integer arithmetic, checked under a database lock on every checkout — which is why tiers here sell out instead of getting slower.

Model id

qwen3.8-27b

Put this in model:. Your plan's ceiling is the longest request it accepts; credits run up to the widest ceiling we serve. Longer requests are refused (HTTP 400), not truncated and not billed.

Quantization

FP8 weights, FP8 (e4m3) KV cache. Every ceiling, price and capacity figure on this page was measured with this configuration; a change to either would be a new model id, never a silent swap.

python
client = OpenAI(base_url="https://api.partitionlabs.ai/v1")stream = client.chat.completions.create(    model="qwen3.8-27b", # one id; your plan sets the ceiling    messages=[{"role": "user", "content": "..."}],    service_tier="flex", # queue order, not speed    stream=True,)

Context ceilings

Every context ceiling is the same model on the same hardware. What you choose is the longest context your subscription reserves — a longer ceiling costs more per share unit because it holds more of the card.

Ceilings on this page: 4k · 8k · 16k · 32k · 64k · 128k · 256k.

Up to 4k context

Requests up to 4k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.

Requests longer than this ceiling are refused on this plan.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$10.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
4k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$40.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
4 concurrent
Flex
4 concurrent
Standard queue
8
Flex queue
16
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
4k tokens
Batch 24h
8 concurrent
Batch 7d
16 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$120.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
12 concurrent
Flex
12 concurrent
Standard queue
24
Flex queue
48
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
4k tokens
Batch 24h
24 concurrent
Batch 7d
48 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$240.00 /mo

The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion

Active
24 concurrent
Flex
24 concurrent
Standard queue
48
Flex queue
96
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
4k tokens
Batch 24h
48 concurrent
Batch 7d
96 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Up to 8k context

Requests up to 8k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.

Requests longer than this ceiling are refused on this plan.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$12.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
8k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$48.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
4 concurrent
Flex
4 concurrent
Standard queue
8
Flex queue
16
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
8k tokens
Batch 24h
8 concurrent
Batch 7d
16 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$144.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
12 concurrent
Flex
12 concurrent
Standard queue
24
Flex queue
48
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
8k tokens
Batch 24h
24 concurrent
Batch 7d
48 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$288.00 /mo

The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion

Active
24 concurrent
Flex
24 concurrent
Standard queue
48
Flex queue
96
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
8k tokens
Batch 24h
48 concurrent
Batch 7d
96 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Up to 16k context

Requests up to 16k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.

Requests longer than this ceiling are refused on this plan.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$14.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
16k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$56.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
4 concurrent
Flex
4 concurrent
Standard queue
8
Flex queue
16
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
16k tokens
Batch 24h
8 concurrent
Batch 7d
16 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$168.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
12 concurrent
Flex
12 concurrent
Standard queue
24
Flex queue
48
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
16k tokens
Batch 24h
24 concurrent
Batch 7d
48 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$336.00 /mo

The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion

Active
24 concurrent
Flex
24 concurrent
Standard queue
48
Flex queue
96
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
16k tokens
Batch 24h
48 concurrent
Batch 7d
96 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Up to 32k context

Requests up to 32k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.

Requests longer than this ceiling are refused on this plan.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$20.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
32k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Duo

Waitlist only

$40.00 /mo

2/24 of an always-on agent's GPU load — at least 2h a day of non-stop processing

Active
2 concurrent
Flex
2 concurrent
Standard queue
4
Flex queue
8
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
32k tokens
Batch 24h
4 concurrent
Batch 7d
8 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$80.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
4 concurrent
Flex
4 concurrent
Standard queue
8
Flex queue
16
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
32k tokens
Batch 24h
8 concurrent
Batch 7d
16 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Pro

Waitlist only

$160.00 /mo

8/24 of an always-on agent's GPU load — at least 8h a day of non-stop processing

Active
8 concurrent
Flex
8 concurrent
Standard queue
16
Flex queue
32
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
32k tokens
Batch 24h
16 concurrent
Batch 7d
32 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Heavy

Waitlist only

$240.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
12 concurrent
Flex
12 concurrent
Standard queue
24
Flex queue
48
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
32k tokens
Batch 24h
24 concurrent
Batch 7d
48 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Max

Waitlist only

$480.00 /mo

The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion

Active
22 concurrent
Flex
24 concurrent
Standard queue
44
Flex queue
96
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
32k tokens
Batch 24h
48 concurrent
Batch 7d
96 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Up to 64k context

Requests up to 64k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.

Requests longer than this ceiling are refused on this plan.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$30.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
64k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$120.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
2 concurrent
Flex
2 concurrent
Standard queue
4
Flex queue
8
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
64k tokens
Batch 24h
4 concurrent
Batch 7d
8 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$360.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
6 concurrent
Flex
6 concurrent
Standard queue
12
Flex queue
24
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
64k tokens
Batch 24h
12 concurrent
Batch 7d
24 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$720.00 /mo

The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion

Active
12 concurrent
Flex
12 concurrent
Standard queue
24
Flex queue
48
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
64k tokens
Batch 24h
24 concurrent
Batch 7d
48 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Up to 128k context

Requests up to 128k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.

Requests longer than this ceiling are refused on this plan.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$50.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
128k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$200.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
128k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$600.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
3 concurrent
Flex
3 concurrent
Standard queue
6
Flex queue
12
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
128k tokens
Batch 24h
6 concurrent
Batch 7d
12 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$1,200.00 /mo

The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion

Active
6 concurrent
Flex
6 concurrent
Standard queue
12
Flex queue
24
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
128k tokens
Batch 24h
12 concurrent
Batch 7d
24 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Up to 256k context

Requests up to 256k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.

Requests longer than this ceiling are refused on this plan.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$90.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
256k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$360.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
256k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$1,080.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
256k tokens
Batch 24h
1 concurrent
Batch 7d
2 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$2,160.00 /mo

The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Reserved context
256k tokens
Batch 24h
1 concurrent
Batch 7d
2 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Hours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.

Beyond one full share. Max is the GPU load of one always-on agent. If you need several, or a pool nothing else runs on, we size it with you rather than from a table.

Talk to us

Pay as you go

Prepaid credits, priced per million tokens on this model. Same model and the same tokens per second as a subscription — what differs is admission.

Credits run up to the widest context ceiling the model serves — one model id, no rung to name. A subscription reserves a ceiling: requests longer than your plan's ceiling are refused rather than run on someone else's reservation.

Pay-as-you-go is best-effort: it fills capacity that reserved tenants are not using, which is most of the time. Under contention a subscriber stays ahead of it — but only inside the share they bought. Past that, their traffic falls below every paying pay-as-you-go request, because the reservation is spent. Nobody is rejected for volume; whoever ranks lower queues.

Priority

queue-jump under contention

Input
$0.80
Cache write
Free
Cached read
$0.10
Output
$6.00

per 1M tokens

Standard

the default

Input
$0.40
Cache write
Free
Cached read
$0.05
Output
$3.00

per 1M tokens

Flex

same model and speed; last in line

Input
$0.30
Cache write
Free
Cached read
$0.0375
Output
$2.25

per 1M tokens

Batch

async; completes within 24h or 7d

Input
$0.20
Cache write
Free
Cached read
$0.025
Output
$1.50

per 1M tokens

No context surcharge at launch.