SLYD
Read first

The 5-step pipeline.

How energy turns into a deployed, financed, offtake-matched cluster, and what SLYD does at every step.

How SLYD works →
Hardware

New and recovered GPU systems.

NVIDIA and AMD systems through documented manufacturer and qualified channel supply, with financing and deployment coordinated on the same platform.

Explore GPU hardware →
Marketplace

Compute, hardware, and power in one book.

Browse available GPU capacity by accelerator, configuration, region, and price, or bring supply to qualified demand.

Open marketplace →
Pre-qualify

Start with an indicative structure.

Tell us deal size, structure, and offtake. Any range is preliminary and subject to underwriting, diligence, and documentation.

Open Configure →
From the blog

GPU market trends and deployment playbooks.

Infrastructure best practices, hardware comparisons, and industry analysis from the SLYD team.

Read the blog →

Library

Enterprise AI Accelerators

GPU Specifications
Database

Specifications for 18 NVIDIA and AMD accelerator records, each showing the manufacturer page it came from and the date it was last checked.

18 records
2 manufacturers
4 architecture families
NVIDIA 14 records
AMD 4 records

Updated

This database lists what NVIDIA and AMD publish for their current AI accelerators: memory, bandwidth, board power, precision performance, and form factor. Each record names the level it describes, because manufacturers publish some figures per accelerator, some only for an eight-GPU platform, and some only for a full rack. Nothing here is converted between those levels or estimated.

The records

How to read these records

  • One accelerator means the figures describe a single GPU.
  • Platform means the figures are totals for a multi-GPU baseboard or module, such as an eight-GPU HGX platform.
  • Rack means the figures are totals for a whole rack-scale system, such as an NVL72.
  • Preliminary specification means the manufacturer states the values are subject to change.
Showing 18 records
NVIDIA Rubin One accelerator

NVIDIA Rubin GPU

Preliminary specification
Memory 288 GB HBM4
Bandwidth 22 TB/s
NVFP4 inference 50 PFLOPS (sparse)
NVFP4 training 35 PFLOPS (dense)
NVLink 3.6 TB/s
Sixth-generation NVLink

NVIDIA footnote: preliminary information, all values are up to and subject to change. The inference figure is sparse; the training figure is dense.

SLYD reading Announced generation. Relevant to roadmap planning, not to a deployment you are specifying now.

NVIDIA Rubin 8-GPU HGX platform

NVIDIA HGX Rubin NVL8

Preliminary specification
Form factor 8x NVIDIA Rubin GPU
Memory 2.3 TB HBM4
Bandwidth 176 TB/s
NVFP4 400 PFLOPS
NVLink Switch 28.8 TB/s
Sixth-generation NVLink Pairs with NVIDIA Vera or x86 CPUs

NVIDIA footnote: preliminary information, all values are up to and subject to change.

SLYD reading Announced 8-GPU server platform for the Rubin generation.

NVIDIA Grace Blackwell Ultra 72-GPU rack-scale system

NVIDIA GB300 NVL72

Shipping
Configuration 72 Blackwell Ultra GPUs, 36 Grace CPUs
GPU memory 20 TB, up to 576 TB/s
Fast memory 37 TB
FP4 Tensor Core 1,440 PFLOPS sparse, 1,080 PFLOPS dense
NVLink 130 TB/s
Fully liquid-cooled rack architecture

NVIDIA publishes this figure as sparse then dense. The dense figure is the one to use for capacity planning unless your workload actually exploits sparsity. These are whole-rack figures across 72 GPUs, not per-accelerator figures.

SLYD reading Rack-scale unit for frontier training and large reasoning-model inference. Specified and sited as a rack, not as servers.

NVIDIA Grace Blackwell 72-GPU rack-scale system

NVIDIA GB200 NVL72

Shipping
Configuration 72 Blackwell GPUs, 36 Grace CPUs
GPU memory 13.4 TB HBM3E, 576 TB/s
CPU memory 17 TB LPDDR5X, 14 TB/s
NVFP4 Tensor Core 1,440 PFLOPS sparse, 720 PFLOPS dense
NVLink 130 TB/s
2,592 Arm Neoverse V2 cores

NVIDIA publishes this figure as sparse then dense. The dense figure is the one to use for capacity planning unless your workload actually exploits sparsity. These are whole-rack figures across 72 GPUs, not per-accelerator figures.

SLYD reading Rack-scale unit for large-model training and high-throughput inference.

NVIDIA Grace Blackwell Superchip module

NVIDIA GB200 Grace Blackwell Superchip

Shipping
Configuration 1 Grace CPU, 2 Blackwell GPUs
Memory 372 GB HBM3E, 16 TB/s
NVLink 3.6 TB/s
Building block of GB200 NVL72

Module-level figures covering two GPUs and one CPU together.

SLYD reading The module NVL72 racks are built from. Useful for understanding rack composition.

NVIDIA Blackwell Ultra 8-GPU HGX platform

NVIDIA HGX B300

Shipping
Form factor 8x NVIDIA Blackwell Ultra SXM
Total memory 2.1 TB
FP4 Tensor Core 144 PFLOPS sparse, 108 PFLOPS dense
FP16/BF16 Tensor Core 36 PFLOPS sparse
FP64 10 TFLOPS
NVLink GPU-to-GPU 1.8 TB/s
Networking 1.6 TB/s
Fifth-generation NVLink Total NVLink bandwidth 14.4 TB/s

NVIDIA publishes this figure as sparse then dense. The dense figure is the one to use for capacity planning unless your workload actually exploits sparsity. NVIDIA publishes these as 8-GPU platform totals. Dividing by eight does not give a supported per-accelerator figure. NVIDIA states HGX B300 is shipping now.

SLYD reading Eight-GPU server platform for training and inference where a full rack-scale system is not the deployment unit.

NVIDIA Blackwell 8-GPU HGX platform

NVIDIA HGX B200

Shipping
Form factor 8x NVIDIA Blackwell SXM
Total memory 1.4 TB
FP4 Tensor Core 144 PFLOPS sparse, 72 PFLOPS dense
FP16/BF16 Tensor Core 36 PFLOPS sparse
FP64 296 TFLOPS
NVLink GPU-to-GPU 1.8 TB/s
Networking 0.8 TB/s
Fifth-generation NVLink Total NVLink bandwidth 14.4 TB/s

NVIDIA publishes this figure as sparse then dense. The dense figure is the one to use for capacity planning unless your workload actually exploits sparsity. NVIDIA publishes these as 8-GPU platform totals. Dividing by eight does not give a supported per-accelerator figure. NVIDIA states HGX B200 is shipping now.

SLYD reading Eight-GPU server platform. Materially stronger FP64 than HGX B300, which matters for mixed AI and HPC sites.

NVIDIA Blackwell, RTX PRO One accelerator

NVIDIA RTX PRO 6000 Blackwell Server Edition

Shipping
Memory 96 GB GDDR7 with ECC
Power 400 to 600 W
Bus PCIe Gen 5 x16
Form factor 4.4 in (H) x 10.5 in (L) dual slot
Thermal Passive
4x DisplayPort 2.1

Passive thermal design: cooling is provided by the host server, so the server must be qualified for it.

SLYD reading Multi-GPU server deployments needing large GDDR7 capacity rather than HBM bandwidth.

NVIDIA Blackwell, RTX PRO One accelerator

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

Shipping
Memory 96 GB GDDR7 with ECC
Power 600 W
Bus PCIe Gen 5 x16
Form factor 5.4 in (H) x 12 in (L) dual slot
Thermal Double flow-through
4x DisplayPort 2.1

SLYD reading Single-GPU workstations. NVIDIA positions this edition for maximum single-GPU throughput.

NVIDIA Blackwell, RTX PRO One accelerator

NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition

Shipping
Memory 96 GB GDDR7 with ECC
Power 300 W
Bus PCIe Gen 5 x16
Form factor 4.4 in (H) x 10.5 in (L) dual slot
Thermal Active
4x DisplayPort 2.1

This is a desktop workstation card, not a laptop or mobile part. NVIDIA publishes it at 300 W in the same dual-slot form factor as the Server Edition.

SLYD reading Dense workstation configurations of up to four GPUs, where total power and thermal budget constrain the build.

NVIDIA Hopper One accelerator

NVIDIA H200 SXM

Shipping
Memory 141 GB HBM3e
Bandwidth 4.8 TB/s
Power Up to 700 W, configurable
Form factor SXM
HGX H200 systems with 4 or 8 GPUs

NVIDIA footnotes these as preliminary specifications that may be subject to change, and marks performance figures as with sparsity. Board power is configurable, so the design figure depends on the OEM system.

SLYD reading Widely deployed generation with a mature software stack. Memory capacity suits models that do not fit an 80 GB part.

NVIDIA Hopper One accelerator

NVIDIA H200 NVL

Shipping
Memory 141 GB HBM3e
Bandwidth 4.8 TB/s
Power Up to 600 W, configurable
Form factor PCIe dual-slot, air-cooled
MGX H200 NVL systems with up to 8 GPUs

Air-cooled PCIe form factor. The same memory and bandwidth as the SXM part at a lower configurable board power, which is what makes it deployable in air-cooled racks.

SLYD reading Sites that need H200 memory capacity but cannot take an SXM baseboard or liquid cooling.

NVIDIA Hopper One accelerator

NVIDIA H100 SXM

Shipping
Memory 80 GB HBM3
Bandwidth 3.35 TB/s
Power Up to 700 W, configurable
Form factor SXM
NVLink 900 GB/s
HGX H100 systems with 4 or 8 GPUs DGX H100 with 8 GPUs

Board power is configurable, so the design figure depends on the OEM system.

SLYD reading The most broadly supported datacenter part in the current installed base. Memory capacity is the usual constraint.

NVIDIA Hopper One accelerator

NVIDIA H100 NVL

Shipping
Memory 94 GB HBM3
Bandwidth 3.9 TB/s
Power 350 to 400 W, configurable
Form factor PCIe dual-slot, air-cooled
NVLink 600 GB/s
Partner and NVIDIA-Certified Systems with 1 to 8 GPUs

More memory and bandwidth than the SXM H100 at roughly half the board power, in an air-cooled PCIe card.

SLYD reading Air-cooled deployments and lower-density racks that cannot support 700 W SXM modules.

AMD CDNA 4 One accelerator

AMD Instinct MI355X

Shipping
Memory 288 GB HBM3E
Bandwidth 8 TB/s peak
Peak MXFP4 matrix 10.1 PFLOPs
Peak FP16 matrix 2.5 PFLOPs dense
Board power 1400 W TBP
Form factor OAM module, PCIe 5.0 x16
Cooling Passive and active
Fourth-generation CDNA Expanded MXFP6 and MXFP4 support

AMD publishes cooling for this part as passive and active. There is no universal liquid-cooling requirement across the Instinct line; the requirement comes from the OEM system.

SLYD reading Largest published memory capacity of the AMD line. Suits large-model serving where capacity per accelerator is the binding constraint.

AMD CDNA 4 One accelerator

AMD Instinct MI350X

Shipping
Memory 288 GB HBM3E
Bandwidth 8 TB/s
Peak MXFP4 matrix 9.2 PFLOPs
Peak FP16 matrix 2.3 PFLOPs
Board power 1000 W TBP
Form factor OAM module
Cooling Passive OAM
Fourth-generation CDNA

Same published memory capacity as MI355X at a lower board power and lower peak matrix performance.

SLYD reading MI355X memory capacity in a 1000 W thermal envelope, for racks that cannot take 1400 W modules.

AMD CDNA One accelerator

AMD Instinct MI325X

Shipping
Memory 256 GB HBM3E
Bandwidth 6 TB/s
Board power 1000 W peak TBP
Form factor OAM module
Cooling Passive OAM
Third-generation CDNA

SLYD reading Prior generation with high memory capacity. Relevant where CDNA 3 software support is already established.

AMD CDNA One accelerator

AMD Instinct MI300X

Shipping
Memory 192 GB HBM3
Bandwidth 5.3 TB/s
Board power 750 W peak TBP
Form factor OAM module
Cooling Passive OAM
Third-generation CDNA

SLYD reading Lowest board power of the Instinct records here, with memory capacity above the 80 GB Hopper part.

Comparison Table

Memory and bandwidth for every record, with the level each figure describes. Precision performance is left to the cards above, because the precisions the two manufacturers publish are not the same and a shared column would invite a comparison the sources do not support.

18 accelerator records from NVIDIA and AMD published specifications, last checked August 18, 2026. The Level column states whether the figures describe one accelerator, a multi-GPU platform, or a full rack.
Record Architecture Level Memory Bandwidth Board power Status
NVIDIA Rubin GPU Rubin One accelerator 288 GB HBM4 22 TB/s Not published Preliminary specification
NVIDIA HGX Rubin NVL8 Rubin 8-GPU HGX platform 2.3 TB HBM4 176 TB/s Not published Preliminary specification
NVIDIA GB300 NVL72 Grace Blackwell Ultra 72-GPU rack-scale system 20 TB, up to 576 TB/s Not published Not published Shipping
NVIDIA GB200 NVL72 Grace Blackwell 72-GPU rack-scale system 13.4 TB HBM3E, 576 TB/s Not published Not published Shipping
NVIDIA GB200 Grace Blackwell Superchip Grace Blackwell Superchip module 372 GB HBM3E, 16 TB/s Not published Not published Shipping
NVIDIA HGX B300 Blackwell Ultra 8-GPU HGX platform 2.1 TB Not published Not published Shipping
NVIDIA HGX B200 Blackwell 8-GPU HGX platform 1.4 TB Not published Not published Shipping
NVIDIA RTX PRO 6000 Blackwell Server Edition Blackwell, RTX PRO One accelerator 96 GB GDDR7 with ECC Not published 400 to 600 W Shipping
NVIDIA RTX PRO 6000 Blackwell Workstation Edition Blackwell, RTX PRO One accelerator 96 GB GDDR7 with ECC Not published 600 W Shipping
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition Blackwell, RTX PRO One accelerator 96 GB GDDR7 with ECC Not published 300 W Shipping
NVIDIA H200 SXM Hopper One accelerator 141 GB HBM3e 4.8 TB/s Up to 700 W, configurable Shipping
NVIDIA H200 NVL Hopper One accelerator 141 GB HBM3e 4.8 TB/s Up to 600 W, configurable Shipping
NVIDIA H100 SXM Hopper One accelerator 80 GB HBM3 3.35 TB/s Up to 700 W, configurable Shipping
NVIDIA H100 NVL Hopper One accelerator 94 GB HBM3 3.9 TB/s 350 to 400 W, configurable Shipping
AMD Instinct MI355X CDNA 4 One accelerator 288 GB HBM3E 8 TB/s peak 1400 W TBP Shipping
AMD Instinct MI350X CDNA 4 One accelerator 288 GB HBM3E 8 TB/s 1000 W TBP Shipping
AMD Instinct MI325X CDNA One accelerator 256 GB HBM3E 6 TB/s 1000 W peak TBP Shipping
AMD Instinct MI300X CDNA One accelerator 192 GB HBM3 5.3 TB/s 750 W peak TBP Shipping

Reading the Database for a Decision

What the specifications here can and cannot tell you about a workload.

Start with memory capacity

The one hard constraint

Memory capacity decides whether a model, its optimizer state, and its KV cache fit at all. Nothing else on this page can compensate for a model that does not fit. Work out your capacity requirement first, then filter the records that clear it.

Treat bandwidth as a ceiling

Not a performance score

Memory bandwidth bounds how fast weights can be streamed, which matters most for memory-bound decode. It does not predict end-to-end throughput or latency on its own: model architecture, quantization, serving software, batching, and topology all move the result.

Match the level to your unit

Accelerator, platform, or rack

If you are buying servers, compare the platform records. If you are siting a rack-scale system, compare the rack records. Comparing a rack figure against a single-accelerator figure is the most common way these tables get misread.

Check power and cooling early

Before the shortlist hardens

Board power here is the accelerator, not the server or the rack. Several parts are configurable, and the design figure comes from the OEM system. Air-cooled PCIe variants exist specifically for sites that cannot take high-power SXM or OAM modules.

Common Questions

Why do some records show figures for a rack or an eight-GPU platform instead of one accelerator?

Because that is the only level at which the manufacturer publishes them. NVIDIA publishes GB300 NVL72 and GB200 NVL72 as 72-GPU rack systems, and HGX B300 and HGX B200 as eight-GPU platforms. Dividing a platform total by eight does not produce a supported per-accelerator figure, so this database labels the scope of every record instead of converting between them.

What does sparse and dense mean in the performance figures?

A sparse figure assumes the workload can exploit structured sparsity in the model weights. A dense figure does not. NVIDIA publishes several precision figures as sparse then dense, and the two differ by up to a factor of two. Use the dense figure for capacity planning unless you have confirmed your model and serving stack actually exploit sparsity.

Why does this database not show prices?

Accelerator pricing depends on configuration, quantity, channel, geography, and the week you ask. A static range with no currency, condition, date, or source is not information a buyer can act on, and SLYD has no governed public price record to publish in its place. Contact SLYD for pricing against a specific configuration.

Does more memory bandwidth mean better performance for my workload?

Not on its own. Memory capacity determines whether a model and its KV cache fit at all, which is a hard constraint. Beyond that, real throughput and latency depend on model architecture, quantization, serving software, batching, and system topology. Treat the figures here as the constraints to design within, not as a performance ranking.

How current are these specifications?

Every record carries the manufacturer page it was read from and the date it was last checked. Those dates are shown on each record and collected in the source register at the foot of this page. Records marked as preliminary carry the manufacturer's own notice that the values are subject to change.

How this database is maintained

Each record is read from the manufacturer's own product page for that product. Reseller listings, comparison articles, and summaries are not used as the source for any figure here.

What is not done

  • No conversion between levels. Where a manufacturer publishes only an eight-GPU platform total, that total is published as a platform figure. It is not divided by eight to state a per-accelerator number.
  • No filling in gaps. A field the manufacturer does not publish on the cited page is left out of that record rather than estimated or marked TBD.
  • No stripping of caveats. Sparsity basis, preliminary-specification notices, and configurable-power notes travel with the figure, because several of these numbers are misleading without them.
  • No price column. Price, availability, and lead time are volatile commercial facts. There is no governed public record behind them, so they are absent rather than stale.

Precision figures

NVIDIA and AMD do not publish the same set of precisions, and the precisions they share are not always measured on the same basis. Precision performance therefore appears on the individual records with its own basis attached, and not as a shared column that would invite a comparison the sources do not support.

Limits of this page

These are published specifications, not measured results on your workload, and not a substitute for the OEM system specification you will actually buy. Board power is the accelerator alone; server input power and rack design load are larger and are covered by the power and cooling calculator. Cooling requirements come from the OEM system, not from the accelerator's cooling field.

Sources and basis

Every specification on this page was read from one of the manufacturer pages below, and is reproduced with that manufacturer's own sparsity, configurability, and preliminary-specification notices attached. Each entry lists the records it supports. Per-record check dates also appear on the individual cards.

  1. AMD Instinct MI300X

    Manufacturer specificationAMD Checked

    https://www.amd.com/en/products/accelerators/instinct/mi300/mi300x.html

  2. AMD Instinct MI325X

    Manufacturer specificationAMD Checked

    https://www.amd.com/en/products/accelerators/instinct/mi300/mi325x.html

  3. AMD Instinct MI350X

    Manufacturer specificationAMD Checked

    https://www.amd.com/en/products/accelerators/instinct/mi350/mi350x.html

  4. AMD Instinct MI355X

    Manufacturer specificationAMD Checked

    https://www.amd.com/en/products/accelerators/instinct/mi350/mi355x.html

  5. NVIDIA GB200 NVL72, NVIDIA GB200 Grace Blackwell Superchip

    Manufacturer specificationNVIDIA Checked

    https://www.nvidia.com/en-us/data-center/gb200-nvl72/

  6. NVIDIA GB300 NVL72

    Manufacturer specificationNVIDIA Checked

    https://www.nvidia.com/en-us/data-center/gb300-nvl72/

  7. NVIDIA H100 SXM, NVIDIA H100 NVL

    Manufacturer specificationNVIDIA Checked

    https://www.nvidia.com/en-us/data-center/h100/

  8. NVIDIA H200 SXM, NVIDIA H200 NVL

    Manufacturer specificationNVIDIA Checked

    https://www.nvidia.com/en-us/data-center/h200/

  9. NVIDIA Rubin GPU, NVIDIA HGX Rubin NVL8, NVIDIA HGX B300, NVIDIA HGX B200

    Manufacturer specificationNVIDIA Checked

    https://www.nvidia.com/en-us/data-center/hgx/

  10. NVIDIA RTX PRO 6000 Blackwell Server Edition, NVIDIA RTX PRO 6000 Blackwell Workstation Edition, NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition

    Manufacturer specificationNVIDIA Checked

    https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-6000-family/

  11. The workload-fit note on each record, and the guidance on reading the database for a decision.

    SLYD analysis

    SLYD's reading of the published specifications. Not a manufacturer claim and not a measured result.

  12. The absence of price, availability, and lead time.

    Reviewed explanation

    These are volatile commercial facts. SLYD publishes no governed public record for them, so this page omits them rather than showing a stale figure.

Need Help Selecting the Right GPU?

Tell us the workload, the memory requirement, and the site constraints, and we will work through the options with you against a specific configuration.

Reconnecting to the server...

Please wait while we restore your connection

An unhandled error has occurred. Reload 🗙