27B multimodal · OpenAI & Anthropic compatible

Qwen 3.8 27B FP8

Qwen 3.8, 27 billion parameters, served in FP8 with image input, streaming and tool calls. One model resident in memory; a subscription picks the context ceiling it is billed at, and nothing else changes between rungs.

Live capacity

Qwen 3.8 27B FP8

Live numbers are unavailable this second. The commitment behind them does not move: committed shares stay at or under capacity as exact integer arithmetic, checked under a database lock on every checkout — which is why tiers here sell out instead of getting slower.

Model ids

One id per context ceiling. Put the one your subscription is billed at in model:; the suffix is the ceiling.

python
client = OpenAI(base_url="https://api.partitionlabs.ai/v1")stream = client.chat.completions.create(    model="qwen3.8-27b-fp8-4k", # the ceiling you subscribed to    messages=[{"role": "user", "content": "..."}],    service_tier="flex", # queue order, not speed    stream=True,)

Context ceilings

Every context ceiling is the same model on the same hardware. What you choose is the longest context your subscription is billed for — nothing else changes between rungs, and a longer ceiling costs more per share unit because it holds more of the card.

Ceilings on this page: 4k · 8k · 16k · 32k · 64k · 128k · 256k.

Up to 4k context

qwen3.8-27b-fp8-4k

Requests up to 4k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$10.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
4k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$40.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
4 concurrent
Flex
4 concurrent
Standard queue
8
Flex queue
16
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
4k tokens
Batch 24h
8 concurrent
Batch 7d
16 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$120.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
12 concurrent
Flex
12 concurrent
Standard queue
24
Flex queue
48
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
4k tokens
Batch 24h
24 concurrent
Batch 7d
48 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$240.00 /mo

The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion

Active
24 concurrent
Flex
24 concurrent
Standard queue
48
Flex queue
96
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
4k tokens
Batch 24h
48 concurrent
Batch 7d
96 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Up to 8k context

qwen3.8-27b-fp8-8k

Requests up to 8k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$12.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
8k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$48.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
4 concurrent
Flex
4 concurrent
Standard queue
8
Flex queue
16
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
8k tokens
Batch 24h
8 concurrent
Batch 7d
16 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$144.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
12 concurrent
Flex
12 concurrent
Standard queue
24
Flex queue
48
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
8k tokens
Batch 24h
24 concurrent
Batch 7d
48 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$288.00 /mo

The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion

Active
24 concurrent
Flex
24 concurrent
Standard queue
48
Flex queue
96
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
8k tokens
Batch 24h
48 concurrent
Batch 7d
96 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Up to 16k context

qwen3.8-27b-fp8-16k

Requests up to 16k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$14.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
16k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$56.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
4 concurrent
Flex
4 concurrent
Standard queue
8
Flex queue
16
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
16k tokens
Batch 24h
8 concurrent
Batch 7d
16 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$168.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
12 concurrent
Flex
12 concurrent
Standard queue
24
Flex queue
48
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
16k tokens
Batch 24h
24 concurrent
Batch 7d
48 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$336.00 /mo

The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion

Active
24 concurrent
Flex
24 concurrent
Standard queue
48
Flex queue
96
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
16k tokens
Batch 24h
48 concurrent
Batch 7d
96 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Up to 32k context

qwen3.8-27b-fp8-32k

Requests up to 32k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$20.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
32k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Duo

Waitlist only

$40.00 /mo

2/24 of an always-on agent's GPU load — at least 2h a day of non-stop processing

Active
2 concurrent
Flex
2 concurrent
Standard queue
4
Flex queue
8
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
32k tokens
Batch 24h
4 concurrent
Batch 7d
8 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$80.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
4 concurrent
Flex
4 concurrent
Standard queue
8
Flex queue
16
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
32k tokens
Batch 24h
8 concurrent
Batch 7d
16 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Pro

Waitlist only

$160.00 /mo

8/24 of an always-on agent's GPU load — at least 8h a day of non-stop processing

Active
8 concurrent
Flex
8 concurrent
Standard queue
16
Flex queue
32
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
32k tokens
Batch 24h
16 concurrent
Batch 7d
32 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Heavy

Waitlist only

$240.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
12 concurrent
Flex
12 concurrent
Standard queue
24
Flex queue
48
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
32k tokens
Batch 24h
24 concurrent
Batch 7d
48 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Max

Waitlist only

$480.00 /mo

The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion

Active
22 concurrent
Flex
24 concurrent
Standard queue
44
Flex queue
96
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
32k tokens
Batch 24h
48 concurrent
Batch 7d
96 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Up to 64k context

qwen3.8-27b-fp8-64k

Requests up to 64k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$30.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
64k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$120.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
2 concurrent
Flex
2 concurrent
Standard queue
4
Flex queue
8
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
64k tokens
Batch 24h
4 concurrent
Batch 7d
8 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$360.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
6 concurrent
Flex
6 concurrent
Standard queue
12
Flex queue
24
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
64k tokens
Batch 24h
12 concurrent
Batch 7d
24 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$720.00 /mo

The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion

Active
12 concurrent
Flex
12 concurrent
Standard queue
24
Flex queue
48
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
64k tokens
Batch 24h
24 concurrent
Batch 7d
48 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Up to 128k context

qwen3.8-27b-fp8-128k

Requests up to 128k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$50.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
128k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$200.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
128k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$600.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
3 concurrent
Flex
3 concurrent
Standard queue
6
Flex queue
12
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
128k tokens
Batch 24h
6 concurrent
Batch 7d
12 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$1,200.00 /mo

The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion

Active
6 concurrent
Flex
6 concurrent
Standard queue
12
Flex queue
24
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
128k tokens
Batch 24h
12 concurrent
Batch 7d
24 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Up to 256k context

qwen3.8-27b-fp8-256k

Requests up to 256k tokens of context. A full share is one always-on request at this length; the longer the context a share has to hold, the more of the card it holds, which is what the price per share unit follows.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$90.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
256k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$360.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
256k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$1,080.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
256k tokens
Batch 24h
1 concurrent
Batch 7d
2 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$2,160.00 /mo

The GPU load of one agent processing non-stop, 24/7 — the most headroom before any demotion

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
256k tokens
Batch 24h
1 concurrent
Batch 7d
2 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Hours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.

Beyond one full share. Max is the GPU load of one always-on agent. If you need several, or a pool nothing else runs on, we size it with you rather than from a table.

Talk to us

Pay as you go

Prepaid credits, priced per million tokens on this model. Same model and the same tokens per second as a subscription — what differs is admission.

Credits run at every context ceiling the model sells. Put the rung you want in model: and the request is admitted up to that ceiling and priced from the model's rate card — the same rule as a subscription.

Pay-as-you-go is best-effort: it fills capacity that reserved tenants are not using, which is most of the time. Under contention a subscriber stays ahead of it — but only inside the share they bought. Past that, their traffic falls below every paying pay-as-you-go request, because the reservation is spent. Nobody is rejected for volume; whoever ranks lower queues.

Priority

queue-jump under contention

Input
$0.80
Cache write
Free
Cached read
$0.10
Output
$6.00

per 1M tokens

Standard

the default

Input
$0.40
Cache write
Free
Cached read
$0.05
Output
$3.00

per 1M tokens

Flex

same model and speed; last in line

Input
$0.30
Cache write
Free
Cached read
$0.0375
Output
$2.25

per 1M tokens

Batch

async; completes within 24h or 7d

Input
$0.20
Cache write
Free
Cached read
$0.025
Output
$1.50

per 1M tokens

No context surcharge at launch.