RAAWRAAWCognitive Systems
Home / Products / Cloud
Infrastructure · GPU Compute

RAAW Cloud

Coming soon

GPU compute, and nothing else. RAAW Cloud rents NVIDIA accelerators by the minute — from a single L4 keeping a model warm to a reserved 8-GPU H100 node for a training run. One rate card, published in full and the same for everyone, with no egress charges, nothing bundled and no minimum commitment.

The lineup

Pick the tier the workload actually needs.

Most GPU spend goes on idle time rather than compute — instances held overnight, or sized for a peak that lasts two hours a day. Choosing the right card is the part you settle at the start, so the lineup is deliberately narrow: one clear choice per class of work, from serving a model to training one from scratch.

// inference tier
Coming soon

L4

The cheapest way to keep a model warm. A single-slot, 72-watt Ada card built for serving rather than training — steady throughput on 7B-class LLMs, embeddings, speech and video pipelines, at a fraction of the power draw of a training GPU.

Memory
24 GB GDDR6
Bandwidth
300 GB/s
Board power
72 W
Best for
Inference · speech
// fine-tune tier
Coming soon

L40S

The workhorse of the fleet. Enough memory and bandwidth to fine-tune mid-size models, run demanding inference and handle rendering or diffusion workloads — without paying for HBM you will not saturate.

Memory
48 GB GDDR6
Bandwidth
864 GB/s
Board power
350 W
Best for
Fine-tuning · diffusion
// training tier
Coming soon

A100 80GB

HBM2e and NVLink, in single cards or full 8-GPU nodes. Still the most cost-effective way to train and fine-tune at scale for most teams — mature drivers, well-understood performance, and far cheaper per hour than current-generation silicon.

Memory
80 GB HBM2e
Bandwidth
2.0 TB/s
Interconnect
NVLink
Form
Single · 8-GPU node
// flagship tier
Coming soon

H100 80GB

Hopper SXM in 8-GPU nodes with NVLink inside the box and InfiniBand between them, for training runs that will not fit anywhere else. Reserved by the week or the month, with the node dedicated to you for its whole term.

Memory
80 GB HBM3
Bandwidth
3.35 TB/s
Interconnect
NVLink · InfiniBand
Form
8-GPU node
// on the roadmap
Roadmap

H200 141GB

141 GB of HBM3e per GPU — room to hold the largest open-weight models in memory without sharding, and the headroom that long-context inference actually needs. Planned for a later phase, once the earlier tiers are running at capacity.

Memory
141 GB HBM3e
Bandwidth
4.8 TB/s
Interconnect
NVLink · InfiniBand
Form
8-GPU node
How it works

One thing, done properly.

// 01

GPU compute only

No general-purpose VMs, no managed databases, no forty services you will never touch. One thing, done properly — which is precisely why we can price it lower.

// 02

Billed by the minute

You pay for the minutes the GPU is attached to you, not a rounded-up hour and not a monthly commitment. Stop the instance and the meter stops with it.

// 03

Data stays where it runs

Your training data and model weights stay in the region the instance runs in — no cross-border transfer, and no copies moved elsewhere for our convenience.

// 04

Ready on first boot

Images that already have CUDA, PyTorch, vLLM and the usual toolchain in place, so an instance is useful the minute it comes up rather than an hour later.

// 05

Persistent storage

NVMe volumes that survive the instance. Detach a GPU overnight to stop paying for it and reattach to the same dataset and checkpoints in the morning.

// 06

Console and API

Launch, snapshot and tear down from the dashboard or straight from your own scripts, with per-project usage visible as it accrues.

Pricing

Published in full when the first region opens.

Every tier will be listed with a plain per-minute rate — no enquiry forms, no quotes, no negotiated discounts that only the largest customers ever see. The same rate applies whether you take one card for an hour or a node for a month, and reserved terms on the training tiers are priced lower again.

Waitlist members get the rate card first, along with launch credits on the inference tiers.

Join the waitlist →

Tell us what you would run.

Startups, labs and studios sizing up GPU capacity can register interest ahead of launch — the workloads we hear about most are the tiers we build out first.

Register interest