GPU catalogue
33 GPU SKUs aggregated from marketplaces — datacenter, workstation, consumer and desktop AI boxes like DGX Spark. Any of them can back a token.
Showing 24 of 33
B300 SXM 288GB
NVIDIA · Blackwell Ultra · 2026
Blackwell Ultra flagship for frontier training and long-context inference. Largest HBM per GPU in the catalogue.
- VRAM
- 288GB
- FP16
- 2.5 PFLOPS
- Rate
- $7.50/hr
GB200 NVL72 (per GPU)
NVIDIA · Blackwell · 2025
Rack-scale Grace Blackwell system billed per GPU. For workloads that need 72 GPUs in one NVLink domain.
- VRAM
- 186GB
- FP16
- 2.25 PFLOPS
- Rate
- $6.90/hr
B200 SXM 180GB
NVIDIA · Blackwell · 2025
Blackwell datacenter GPU with 180GB HBM3e. Suited for training and serving the largest open-weight models on a single node.
- VRAM
- 180GB
- FP16
- 2.25 PFLOPS
- Rate
- $5.50/hr
Instinct MI355X 288GB
AMD · CDNA 4 · 2025
AMD's CDNA 4 flagship with 288GB HBM3e and FP4/FP6 support. ROCm-first stack for large-model serving.
- VRAM
- 288GB
- FP16
- 2.3 PFLOPS
- Rate
- $4.50/hr
H200 SXM 141GB
NVIDIA · Hopper · 2024
Hopper with 141GB HBM3e and 4.8TB/s bandwidth. The sweet spot for serving 70B–400B models with long context.
- VRAM
- 141GB
- FP16
- 989 TFLOPS
- Rate
- $3.60/hr
Instinct MI325X 256GB
AMD · CDNA 3 · 2024
CDNA 3 with 256GB HBM3e. Memory-heavy inference for models that do not fit on 141GB.
- VRAM
- 256GB
- FP16
- 1.31 PFLOPS
- Rate
- $3.20/hr
H100 NVL 94GB
NVIDIA · Hopper · 2023
PCIe H100 with 94GB and an NVLink bridge between pairs. Built for LLM inference in standard servers.
- VRAM
- 94GB
- FP16
- 835 TFLOPS
- Rate
- $2.80/hr
H100 SXM 80GB
NVIDIA · Hopper · 2022
The workhorse of modern AI. SXM form factor with NVLink for multi-GPU training and high-throughput inference.
- VRAM
- 80GB
- FP16
- 989 TFLOPS
- Rate
- $2.69/hr
Instinct MI300X 192GB
AMD · CDNA 3 · 2023
192GB HBM3 per GPU — runs a 70B model in BF16 on a single card. Strong vLLM and SGLang support.
- VRAM
- 192GB
- FP16
- 1.31 PFLOPS
- Rate
- $2.40/hr
H100 PCIe 80GB
NVIDIA · Hopper · 2022
PCIe H100 for single-GPU fine-tuning and inference where NVLink is not required.
- VRAM
- 80GB
- FP16
- 756 TFLOPS
- Rate
- $2.20/hr
Gaudi 3 128GB
Intel · Gaudi · 2024
Intel's training and inference accelerator with 128GB HBM2e and built-in Ethernet scale-out. Runs PyTorch via Habana SynapseAI.
- VRAM
- 128GB
- FP16
- 1.83 PFLOPS
- Rate
- $2.00/hr
GH200 Grace Hopper 96GB
NVIDIA · Hopper · 2023
Hopper GPU fused to a 72-core Grace CPU with 480GB LPDDR5X. Huge unified memory for inference that spills past HBM.
- VRAM
- 96GB
- FP16
- 989 TFLOPS
- Rate
- $1.85/hr
A100 SXM 80GB
NVIDIA · Ampere · 2020
Ampere datacenter GPU. Still the best value for training mid-size models and serving 7B–70B models.
- VRAM
- 80GB
- FP16
- 312 TFLOPS
- Rate
- $1.60/hr
Instinct MI250X 128GB
AMD · CDNA 2 · 2021
Previous-gen CDNA 2 accelerator with 128GB. Plentiful and cheap on HPC-oriented clouds.
- VRAM
- 128GB
- FP16
- 383 TFLOPS
- Rate
- $1.50/hr
A100 PCIe 80GB
NVIDIA · Ampere · 2021
PCIe A100 for standard servers. Same compute as SXM with lower interconnect bandwidth.
- VRAM
- 80GB
- FP16
- 312 TFLOPS
- Rate
- $1.40/hr
A100 PCIe 40GB
NVIDIA · Ampere · 2020
The original A100. 40GB is enough for most fine-tuning runs under 13B parameters.
- VRAM
- 40GB
- FP16
- 312 TFLOPS
- Rate
- $1.10/hr
L40S 48GB
NVIDIA · Ada Lovelace · 2023
Ada datacenter GPU for inference, rendering and video. FP8 support and 48GB make it a strong single-GPU server.
- VRAM
- 48GB
- FP16
- 362 TFLOPS
- Rate
- $0.90/hr
L40 48GB
NVIDIA · Ada Lovelace · 2022
Visual-computing Ada GPU. Good for diffusion models, rendering and moderate LLM inference.
- VRAM
- 48GB
- FP16
- 181 TFLOPS
- Rate
- $0.80/hr
A40 48GB
NVIDIA · Ampere · 2020
Ampere visual-computing GPU with 48GB. Cheap capacity for inference that needs memory more than speed.
- VRAM
- 48GB
- FP16
- 150 TFLOPS
- Rate
- $0.40/hr
L4 24GB
NVIDIA · Ada Lovelace · 2023
Low-power Ada inference card. Ideal for small models, video transcoding and embeddings at scale.
- VRAM
- 24GB
- FP16
- 121 TFLOPS
- Rate
- $0.40/hr
V100 SXM 16GB
NVIDIA · Volta · 2017
Legacy Volta GPU. Cheap tensor cores for classic deep learning and small fine-tunes.
- VRAM
- 16GB
- FP16
- 125 TFLOPS
- Rate
- $0.20/hr
T4 16GB
NVIDIA · Turing · 2018
70W inference card. The cheapest CUDA GPU on most clouds — fine for embeddings and tiny models.
- VRAM
- 16GB
- FP16
- 65 TFLOPS
- Rate
- $0.15/hr
DGX Station (GB300) 784GB
NVIDIA · Blackwell Ultra · 2025
Deskside Grace Blackwell Ultra with 784GB coherent memory. A single box that fits trillion-parameter inference.
- VRAM
- 784GB
- FP16
- 2.5 PFLOPS
- Rate
- $9.00/hr
DGX Spark (GB10) 128GB
NVIDIA · Blackwell · 2025
NVIDIA's desktop AI supercomputer: Grace Blackwell GB10 with 128GB unified memory. Hosted by the community — runs 200B-parameter models locally.
- VRAM
- 128GB
- FP16
- 125 TFLOPS
- Rate
- $0.90/hr