RAAW Cloud
Coming soonGPU compute, and nothing else. RAAW Cloud rents NVIDIA accelerators by the minute — from a single L4 keeping a model warm to a reserved 8-GPU H100 node for a training run. One rate card, published in full and the same for everyone, with no egress charges, nothing bundled and no minimum commitment.
Pick the tier the workload actually needs.
Most GPU spend goes on idle time rather than compute — instances held overnight, or sized for a peak that lasts two hours a day. Choosing the right card is the part you settle at the start, so the lineup is deliberately narrow: one clear choice per class of work, from serving a model to training one from scratch.
L4
The cheapest way to keep a model warm. A single-slot, 72-watt Ada card built for serving rather than training — steady throughput on 7B-class LLMs, embeddings, speech and video pipelines, at a fraction of the power draw of a training GPU.
L40S
The workhorse of the fleet. Enough memory and bandwidth to fine-tune mid-size models, run demanding inference and handle rendering or diffusion workloads — without paying for HBM you will not saturate.
A100 80GB
HBM2e and NVLink, in single cards or full 8-GPU nodes. Still the most cost-effective way to train and fine-tune at scale for most teams — mature drivers, well-understood performance, and far cheaper per hour than current-generation silicon.
H100 80GB
Hopper SXM in 8-GPU nodes with NVLink inside the box and InfiniBand between them, for training runs that will not fit anywhere else. Reserved by the week or the month, with the node dedicated to you for its whole term.
H200 141GB
141 GB of HBM3e per GPU — room to hold the largest open-weight models in memory without sharding, and the headroom that long-context inference actually needs. Planned for a later phase, once the earlier tiers are running at capacity.
One thing, done properly.
GPU compute only
No general-purpose VMs, no managed databases, no forty services you will never touch. One thing, done properly — which is precisely why we can price it lower.
Billed by the minute
You pay for the minutes the GPU is attached to you, not a rounded-up hour and not a monthly commitment. Stop the instance and the meter stops with it.
Data stays where it runs
Your training data and model weights stay in the region the instance runs in — no cross-border transfer, and no copies moved elsewhere for our convenience.
Ready on first boot
Images that already have CUDA, PyTorch, vLLM and the usual toolchain in place, so an instance is useful the minute it comes up rather than an hour later.
Persistent storage
NVMe volumes that survive the instance. Detach a GPU overnight to stop paying for it and reattach to the same dataset and checkpoints in the morning.
Console and API
Launch, snapshot and tear down from the dashboard or straight from your own scripts, with per-project usage visible as it accrues.
Published in full when the first region opens.
Every tier will be listed with a plain per-minute rate — no enquiry forms, no quotes, no negotiated discounts that only the largest customers ever see. The same rate applies whether you take one card for an hour or a node for a month, and reserved terms on the training tiers are priced lower again.
Waitlist members get the rate card first, along with launch credits on the inference tiers.
Join the waitlist →Tell us what you would run.
Startups, labs and studios sizing up GPU capacity can register interest ahead of launch — the workloads we hear about most are the tiers we build out first.
Register interest