- 180 GB HBM3e
- ~18 PFLOPS FP4
- NVLink + high-bandwidth interconnect
Compute for every
AI workload.
A full NVIDIA GPU lineup on one platform — H100, H200, B200, L40S and A100. Pick the abstraction level your team needs.
Pick the right muscle for the job
From single-GPU dev instances to multi-thousand-GPU training clusters. Every node is NVLink-connected with high-bandwidth interconnect. Specs below are common reference values — confirm exact configuration at booking.
- 141 GB HBM3e
- ~4 PFLOPS FP8
- NVLink + high-bandwidth interconnect
- 80 GB HBM3
- ~4 PFLOPS FP8
- NVLink + high-bandwidth interconnect
- 48 GB GDDR6
- ~1466 TFLOPS FP8
- NVLink + high-bandwidth interconnect
- 80 GB HBM2e
- ~624 TFLOPS FP16
- NVLink + high-bandwidth interconnect
* Tensor Core performance figures include sparsity acceleration. Actual performance varies by workload, model, and deployment configuration.
Three ways to consume
From raw metal to managed inference — choose the operating model that fits your team.
Bare-Metal
NVIDIA H100 / H200 / B200, up to 8 GPUs per node. NVLink + high-bandwidth interconnect. No virtualization overhead.
Request a quoteKubernetes
On-demand & reserved GPU tiers with managed K8s, autoscaling and portable workloads.
Request a quoteInference
Low-latency serving from regional edge, with autoscaling endpoints for LLMs and vision models.
Request a quote