GPU Lineup

Compute for every
AI workload.

A full NVIDIA GPU lineup on one platform — H100, H200, B200, L40S and A100. Pick the abstraction level your team needs.

Pick the right muscle for the job

From single-GPU dev instances to multi-thousand-GPU training clusters. Every node is NVLink-connected with high-bandwidth interconnect. Specs below are common reference values — confirm exact configuration at booking.

NVIDIA B200
  • 180 GB HBM3e
  • ~18 PFLOPS FP4
  • NVLink + high-bandwidth interconnect
Request a quote
NVIDIA H200
  • 141 GB HBM3e
  • ~4 PFLOPS FP8
  • NVLink + high-bandwidth interconnect
Request a quote
NVIDIA H100
  • 80 GB HBM3
  • ~4 PFLOPS FP8
  • NVLink + high-bandwidth interconnect
Request a quote
NVIDIA L40S
  • 48 GB GDDR6
  • ~1466 TFLOPS FP8
  • NVLink + high-bandwidth interconnect
Request a quote
NVIDIA A100
  • 80 GB HBM2e
  • ~624 TFLOPS FP16
  • NVLink + high-bandwidth interconnect
Request a quote

* Tensor Core performance figures include sparsity acceleration. Actual performance varies by workload, model, and deployment configuration.

Three ways to consume

From raw metal to managed inference — choose the operating model that fits your team.

🖥️

Bare-Metal

NVIDIA H100 / H200 / B200, up to 8 GPUs per node. NVLink + high-bandwidth interconnect. No virtualization overhead.

Request a quote
⚙️

Kubernetes

On-demand & reserved GPU tiers with managed K8s, autoscaling and portable workloads.

Request a quote
🚀

Inference

Low-latency serving from regional edge, with autoscaling endpoints for LLMs and vision models.

Request a quote

Ready to size your workload?

Contact our team for a tailored quote.