Pricing
Subscriptions buy a guaranteed share of a pool. Pay-as-you-go credits fill the slack, and both run the same model at the same speed — service tiers change queue order, never tokens per second.
Subscriptions
Every pool runs one model, resident in memory. A subscription reserves share on one pool, not on the fleet, so what you buy is backed by capacity counted against those weights rather than by an average across all of them.
Qwen 3.8 27B FP8
qwen3.8-27b-fp8-4k
4k context pool. 4k context. One full share is one always-on 4k request; a share unit is priced at $10/month because F(4k) is what the card actually holds at that length.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$10.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 4k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$40.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 4 concurrent
- Flex
- 4 concurrent
- Standard queue
- 8
- Flex queue
- 16
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 4k tokens
- Batch 24h
- 8 concurrent
- Batch 7d
- 16 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$120.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 12 concurrent
- Flex
- 12 concurrent
- Standard queue
- 24
- Flex queue
- 48
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 4k tokens
- Batch 24h
- 24 concurrent
- Batch 7d
- 48 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$240.00 /mo
The GPU load of one agent processing non-stop, 24/7 — top subscriber priority
- Active
- 24 concurrent
- Flex
- 24 concurrent
- Standard queue
- 48
- Flex queue
- 96
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 4k tokens
- Batch 24h
- 48 concurrent
- Batch 7d
- 96 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.
Qwen 3.8 27B FP8
qwen3.8-27b-fp8-8k
8k context pool. 8k context. One full share is one always-on 8k request; a share unit is priced at $12/month because F(8k) is what the card actually holds at that length.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$12.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 8k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$48.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 4 concurrent
- Flex
- 4 concurrent
- Standard queue
- 8
- Flex queue
- 16
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 8k tokens
- Batch 24h
- 8 concurrent
- Batch 7d
- 16 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$144.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 12 concurrent
- Flex
- 12 concurrent
- Standard queue
- 24
- Flex queue
- 48
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 8k tokens
- Batch 24h
- 24 concurrent
- Batch 7d
- 48 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$288.00 /mo
The GPU load of one agent processing non-stop, 24/7 — top subscriber priority
- Active
- 24 concurrent
- Flex
- 24 concurrent
- Standard queue
- 48
- Flex queue
- 96
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 8k tokens
- Batch 24h
- 48 concurrent
- Batch 7d
- 96 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.
Qwen 3.8 27B FP8
qwen3.8-27b-fp8-16k
16k context pool. 16k context. One full share is one always-on 16k request; a share unit is priced at $14/month because F(16k) is what the card actually holds at that length.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$14.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 16k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$56.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 4 concurrent
- Flex
- 4 concurrent
- Standard queue
- 8
- Flex queue
- 16
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 16k tokens
- Batch 24h
- 8 concurrent
- Batch 7d
- 16 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$168.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 12 concurrent
- Flex
- 12 concurrent
- Standard queue
- 24
- Flex queue
- 48
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 16k tokens
- Batch 24h
- 24 concurrent
- Batch 7d
- 48 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$336.00 /mo
The GPU load of one agent processing non-stop, 24/7 — top subscriber priority
- Active
- 24 concurrent
- Flex
- 24 concurrent
- Standard queue
- 48
- Flex queue
- 96
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 16k tokens
- Batch 24h
- 48 concurrent
- Batch 7d
- 96 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.
Qwen 3.8 27B FP8
qwen3.8-27b-fp8
General pool. 32k context. One full share is one always-on 32k request; a share unit is priced at $20/month because F(32k) is what the card actually holds at that length.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$20.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 32k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDuo
Waitlist only$40.00 /mo
2/24 of an always-on agent's GPU load — at least 2h a day of non-stop processing
- Active
- 2 concurrent
- Flex
- 2 concurrent
- Standard queue
- 4
- Flex queue
- 8
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 32k tokens
- Batch 24h
- 4 concurrent
- Batch 7d
- 8 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$80.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 4 concurrent
- Flex
- 4 concurrent
- Standard queue
- 8
- Flex queue
- 16
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 32k tokens
- Batch 24h
- 8 concurrent
- Batch 7d
- 16 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistPro
Waitlist only$160.00 /mo
8/24 of an always-on agent's GPU load — at least 8h a day of non-stop processing
- Active
- 8 concurrent
- Flex
- 8 concurrent
- Standard queue
- 16
- Flex queue
- 32
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 32k tokens
- Batch 24h
- 16 concurrent
- Batch 7d
- 32 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHeavy
Waitlist only$240.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 12 concurrent
- Flex
- 12 concurrent
- Standard queue
- 24
- Flex queue
- 48
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 32k tokens
- Batch 24h
- 24 concurrent
- Batch 7d
- 48 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistMax
Waitlist only$480.00 /mo
The GPU load of one agent processing non-stop, 24/7 — top subscriber priority
- Active
- 22 concurrent
- Flex
- 24 concurrent
- Standard queue
- 44
- Flex queue
- 96
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 32k tokens
- Batch 24h
- 48 concurrent
- Batch 7d
- 96 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.
Qwen 3.8 27B FP8
qwen3.8-27b-fp8-64k
64k context pool. 64k context. One full share is one always-on 64k request; a share unit is priced at $30/month because F(64k) is what the card actually holds at that length.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$30.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 64k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$120.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 2 concurrent
- Flex
- 2 concurrent
- Standard queue
- 4
- Flex queue
- 8
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 64k tokens
- Batch 24h
- 4 concurrent
- Batch 7d
- 8 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$360.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 6 concurrent
- Flex
- 6 concurrent
- Standard queue
- 12
- Flex queue
- 24
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 64k tokens
- Batch 24h
- 12 concurrent
- Batch 7d
- 24 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$720.00 /mo
The GPU load of one agent processing non-stop, 24/7 — top subscriber priority
- Active
- 12 concurrent
- Flex
- 12 concurrent
- Standard queue
- 24
- Flex queue
- 48
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 64k tokens
- Batch 24h
- 24 concurrent
- Batch 7d
- 48 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.
Qwen 3.8 27B FP8
qwen3.8-27b-fp8-128k
128k context pool. 128k context. One full share is one always-on 128k request; a share unit is priced at $50/month because F(128k) is what the card actually holds at that length.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$50.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 128k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$200.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 128k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$600.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 3 concurrent
- Flex
- 3 concurrent
- Standard queue
- 6
- Flex queue
- 12
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 128k tokens
- Batch 24h
- 6 concurrent
- Batch 7d
- 12 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$1,200.00 /mo
The GPU load of one agent processing non-stop, 24/7 — top subscriber priority
- Active
- 6 concurrent
- Flex
- 6 concurrent
- Standard queue
- 12
- Flex queue
- 24
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 128k tokens
- Batch 24h
- 12 concurrent
- Batch 7d
- 24 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.
Qwen 3.8 27B FP8
qwen3.8-27b-fp8-256k
256k context pool. 256k context. One full share is one always-on 256k request; a share unit is priced at $90/month because F(256k) is what the card actually holds at that length.
Live capacity is unavailable right now — we never sell more shares than we can serve.
Solo
Waitlist only$90.00 /mo
1/24 of an always-on agent's GPU load — at least 1h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 256k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistDev
Waitlist only$360.00 /mo
4/24 of an always-on agent's GPU load — at least 4h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 256k tokens
- Batch 24h
- 2 concurrent
- Batch 7d
- 4 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHalf share
Waitlist only$1,080.00 /mo
12/24 of an always-on agent's GPU load — at least 12h a day of non-stop processing
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 256k tokens
- Batch 24h
- 1 concurrent
- Batch 7d
- 2 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistFull share
Waitlist only$2,160.00 /mo
The GPU load of one agent processing non-stop, 24/7 — top subscriber priority
- Active
- 1 concurrent
- Flex
- 1 concurrent
- Standard queue
- 2
- Flex queue
- 4
- Max standard wait
- 60s, then 429
- Max flex wait
- 240s, then 429
- Max context
- 256k tokens
- Batch 24h
- 1 concurrent
- Batch 7d
- 2 concurrent
Opening capacity in batches — email to be notified
Email us to join the waitlistHours are a floor, counted against a fully sold pool. Reserved capacity is guaranteed under contention; additional usage is available on a best-effort basis, so a quiet pool potentially gives you more than your share. Interactive work is bursty rather than continuous — an hour of non-stop processing typically covers two to four hours of interactive hands-on development.
Beyond one full share. Max is the GPU load of one always-on agent. If you need several, or a pool nothing else runs on, we size it with you rather than from a table.
Talk to usPay as you go
Prepaid credits, priced per million tokens. Same model and the same tokens per second as a subscription — what differs is admission.
Pay-as-you-go is best-effort: it fills capacity that reserved tenants are not using, which is most of the time. Under contention a subscriber stays ahead of it — but only inside the share they bought. Past that, their traffic falls below every paying pay-as-you-go request, because the reservation is spent. Nobody is rejected for volume; whoever ranks lower queues.
Priority
queue-jump under contention
- Input
- $0.80
- Cache write
- Free
- Cached read
- $0.10
- Output
- $6.00
per 1M tokens
Standard
the default
- Input
- $0.40
- Cache write
- Free
- Cached read
- $0.05
- Output
- $3.00
per 1M tokens
Flex
same model and speed; last in line
- Input
- $0.30
- Cache write
- Free
- Cached read
- $0.0375
- Output
- $2.25
per 1M tokens
Batch
async; completes within 24h or 7d
- Input
- $0.20
- Cache write
- Free
- Cached read
- $0.025
- Output
- $1.50
per 1M tokens
No context surcharge at launch.
Batch finishes inside the window you pick — 24 hours or 7 days, same price. We refuse work we cannot schedule in time rather than accept it and run late; anything still unfinished at the deadline expires and is not billed. Results stay available for 30 days after completion, or less if you ask. Files up to 512 MiB, batches up to 100,000 requests or 256 MiB; storage is included with every plan.
Credits are bought in packages starting at $5.00. Credits expire 12 months after purchase.