Pricing

Subscriptions buy a guaranteed share of a pool. Pay-as-you-go credits fill the slack, and both run the same model at the same speed — service tiers change queue order, never tokens per second.

Subscriptions

Every pool runs one model, resident in memory. A subscription reserves share on one pool, not on the fleet, so what you buy is backed by capacity counted against those weights rather than by an average across all of them.

Qwen 3.8 27B FP8

qwen3.8-27b-fp8-4k

4k context pool. 4k context. One full share is one always-on 4k request; a share unit is priced at $10/month because F(4k) is what the card actually holds at that length.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$10.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
4k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$40.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
4 concurrent
Flex
4 concurrent
Standard queue
8
Flex queue
16
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
4k tokens
Batch 24h
8 concurrent
Batch 7d
16 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$120.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
12 concurrent
Flex
12 concurrent
Standard queue
24
Flex queue
48
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
4k tokens
Batch 24h
24 concurrent
Batch 7d
48 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$240.00 /mo

The GPU load of one agent processing non-stop, 24/7 — top subscriber priority

Active
24 concurrent
Flex
24 concurrent
Standard queue
48
Flex queue
96
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
4k tokens
Batch 24h
48 concurrent
Batch 7d
96 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Hours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.

Qwen 3.8 27B FP8

qwen3.8-27b-fp8-8k

8k context pool. 8k context. One full share is one always-on 8k request; a share unit is priced at $12/month because F(8k) is what the card actually holds at that length.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$12.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
8k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$48.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
4 concurrent
Flex
4 concurrent
Standard queue
8
Flex queue
16
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
8k tokens
Batch 24h
8 concurrent
Batch 7d
16 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$144.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
12 concurrent
Flex
12 concurrent
Standard queue
24
Flex queue
48
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
8k tokens
Batch 24h
24 concurrent
Batch 7d
48 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$288.00 /mo

The GPU load of one agent processing non-stop, 24/7 — top subscriber priority

Active
24 concurrent
Flex
24 concurrent
Standard queue
48
Flex queue
96
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
8k tokens
Batch 24h
48 concurrent
Batch 7d
96 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Hours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.

Qwen 3.8 27B FP8

qwen3.8-27b-fp8-16k

16k context pool. 16k context. One full share is one always-on 16k request; a share unit is priced at $14/month because F(16k) is what the card actually holds at that length.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$14.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
16k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$56.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
4 concurrent
Flex
4 concurrent
Standard queue
8
Flex queue
16
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
16k tokens
Batch 24h
8 concurrent
Batch 7d
16 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$168.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
12 concurrent
Flex
12 concurrent
Standard queue
24
Flex queue
48
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
16k tokens
Batch 24h
24 concurrent
Batch 7d
48 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$336.00 /mo

The GPU load of one agent processing non-stop, 24/7 — top subscriber priority

Active
24 concurrent
Flex
24 concurrent
Standard queue
48
Flex queue
96
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
16k tokens
Batch 24h
48 concurrent
Batch 7d
96 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Hours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.

Qwen 3.8 27B FP8

qwen3.8-27b-fp8

General pool. 32k context. One full share is one always-on 32k request; a share unit is priced at $20/month because F(32k) is what the card actually holds at that length.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$20.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
32k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Duo

Waitlist only

$40.00 /mo

2/24 of an always-on agent's GPU load — at least 2h a day of non-stop processing

Active
2 concurrent
Flex
2 concurrent
Standard queue
4
Flex queue
8
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
32k tokens
Batch 24h
4 concurrent
Batch 7d
8 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$80.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
4 concurrent
Flex
4 concurrent
Standard queue
8
Flex queue
16
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
32k tokens
Batch 24h
8 concurrent
Batch 7d
16 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Pro

Waitlist only

$160.00 /mo

8/24 of an always-on agent's GPU load — at least 8h a day of non-stop processing

Active
8 concurrent
Flex
8 concurrent
Standard queue
16
Flex queue
32
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
32k tokens
Batch 24h
16 concurrent
Batch 7d
32 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Heavy

Waitlist only

$240.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
12 concurrent
Flex
12 concurrent
Standard queue
24
Flex queue
48
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
32k tokens
Batch 24h
24 concurrent
Batch 7d
48 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Max

Waitlist only

$480.00 /mo

The GPU load of one agent processing non-stop, 24/7 — top subscriber priority

Active
22 concurrent
Flex
24 concurrent
Standard queue
44
Flex queue
96
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
32k tokens
Batch 24h
48 concurrent
Batch 7d
96 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Hours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.

Qwen 3.8 27B FP8

qwen3.8-27b-fp8-64k

64k context pool. 64k context. One full share is one always-on 64k request; a share unit is priced at $30/month because F(64k) is what the card actually holds at that length.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$30.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
64k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$120.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
2 concurrent
Flex
2 concurrent
Standard queue
4
Flex queue
8
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
64k tokens
Batch 24h
4 concurrent
Batch 7d
8 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$360.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
6 concurrent
Flex
6 concurrent
Standard queue
12
Flex queue
24
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
64k tokens
Batch 24h
12 concurrent
Batch 7d
24 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$720.00 /mo

The GPU load of one agent processing non-stop, 24/7 — top subscriber priority

Active
12 concurrent
Flex
12 concurrent
Standard queue
24
Flex queue
48
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
64k tokens
Batch 24h
24 concurrent
Batch 7d
48 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Hours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.

Qwen 3.8 27B FP8

qwen3.8-27b-fp8-128k

128k context pool. 128k context. One full share is one always-on 128k request; a share unit is priced at $50/month because F(128k) is what the card actually holds at that length.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$50.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
128k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$200.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
128k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$600.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
3 concurrent
Flex
3 concurrent
Standard queue
6
Flex queue
12
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
128k tokens
Batch 24h
6 concurrent
Batch 7d
12 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$1,200.00 /mo

The GPU load of one agent processing non-stop, 24/7 — top subscriber priority

Active
6 concurrent
Flex
6 concurrent
Standard queue
12
Flex queue
24
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
128k tokens
Batch 24h
12 concurrent
Batch 7d
24 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Hours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.

Qwen 3.8 27B FP8

qwen3.8-27b-fp8-256k

256k context pool. 256k context. One full share is one always-on 256k request; a share unit is priced at $90/month because F(256k) is what the card actually holds at that length.

Live capacity is unavailable right now — we never sell more shares than we can serve.

Solo

Waitlist only

$90.00 /mo

1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
256k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Dev

Waitlist only

$360.00 /mo

4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
256k tokens
Batch 24h
2 concurrent
Batch 7d
4 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Half share

Waitlist only

$1,080.00 /mo

12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
256k tokens
Batch 24h
1 concurrent
Batch 7d
2 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Full share

Waitlist only

$2,160.00 /mo

The GPU load of one agent processing non-stop, 24/7 — top subscriber priority

Active
1 concurrent
Flex
1 concurrent
Standard queue
2
Flex queue
4
Max standard wait
60s, then 429
Max flex wait
240s, then 429
Max context
256k tokens
Batch 24h
1 concurrent
Batch 7d
2 concurrent

Opening capacity in batches — email to be notified

Email us to join the waitlist

Hours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.

Beyond one full share. Max is the GPU load of one always-on agent. If you need several, or a pool nothing else runs on, we size it with you rather than from a table.

Talk to us

Pay as you go

Prepaid credits, priced per million tokens. Same model and the same tokens per second as a subscription — what differs is admission.

Pay-as-you-go is best-effort: it fills capacity that reserved tenants are not using, which is most of the time. Under contention a subscriber stays ahead of it — but only inside the share they bought. Past that, their traffic falls below every paying pay-as-you-go request, because the reservation is spent. Nobody is rejected for volume; whoever ranks lower queues.

Priority

queue-jump under contention

Input
$0.80
Cache write
Free
Cached read
$0.10
Output
$6.00

per 1M tokens

Standard

the default

Input
$0.40
Cache write
Free
Cached read
$0.05
Output
$3.00

per 1M tokens

Flex

same model and speed; last in line

Input
$0.30
Cache write
Free
Cached read
$0.0375
Output
$2.25

per 1M tokens

Batch

async; completes within 24h or 7d

Input
$0.20
Cache write
Free
Cached read
$0.025
Output
$1.50

per 1M tokens

No context surcharge at launch.

Batch finishes inside the window you pick — 24 hours or 7 days, same price. We refuse work we cannot schedule in time rather than accept it and run late; anything still unfinished at the deadline expires and is not billed. Results stay available for 30 days after completion, or less if you ask. Files up to 512 MiB, batches up to 100,000 requests or 256 MiB; storage is included with every plan.

Credits are bought in packages starting at $5.00. Credits expire 12 months after purchase.