We take cryptocurrency — so the coins you hold can buy real hardware.Pay in crypto, spend it on real hardware. Free insured shipping worldwide over $399. Plain, unbranded boxes, sent from Germany.Free insured shipping over $399. Your rate is locked for 30 minutes once checkout opens.Rate locked 30 minutes. Every item sealed, serial-checked and covered by a 24-month warranty.Sealed, 24-month warranty.

Saved items Sign in

HPE

HPE Q0E21A NVIDIA Tesla P100 16 GB HBM2 workstation graphics card

SKU CH-QFN9-9XKF7

A 16 GB HBM2 computational accelerator with 3584 CUDA cores and 732 GB/s memory bandwidth for HPC and deep learning workloads

  • InterfacePCI Express 3.0 x16
  • Chipset ManufacturerNVIDIA
  • GPUTesla P100
  • Core Clock1190 MHz
  • CUDA Cores3584
  • Memory Clock715 MHz
  • Pascal architecture drives 9.3 TFLOPS single-precision compute
  • 16 GB HBM2 memory with 732 GB/s bandwidth
  • 3584 CUDA cores accelerate parallel workloads
  • 7 TFLOPS double-precision for HPC tasks
  • PCI Express 3.0 x16 interface fits dual-slot workstations

HPE Q0E21A Tesla P100 16 GB HBM2 accelerator

The HPE Q0E21A houses an NVIDIA Tesla P100 GPU built on Pascal architecture with 3584 CUDA cores and 16 GB of HBM2 memory across a 4096-bit interface. It targets high-performance computing and deep-learning workloads that demand high double-precision throughput and memory bandwidth of 732 GB/s. The card draws up to 250 W and occupies a dual-slot footprint measuring 267 mm long. It is an OEM version supplied without retail packaging.

PCI Express 3.0 x16 dual-slot installation

The board installs in a PCI Express 3.0 x16 slot and requires a chassis that can accommodate a 267 mm dual-slot card and deliver 250 W through the appropriate power connectors. Its HBM2 memory is stacked on the GPU package using CoWoS technology, eliminating separate VRAM chips and freeing PCB space for dense server layouts. Builders must verify that their system provides sufficient airflow and power headroom for sustained computational tasks.

Sustained double-precision output of 4.7 teraFLOPS

Under continuous load the Tesla P100 maintains a core clock of 1190 MHz with memory at 715 MHz, delivering 4.7 teraFLOPS double-precision and 9.3 teraFLOPS single-precision performance. The 4096-bit HBM2 interface sustains 732 GB/s bandwidth, reducing bottlenecks in memory-bound HPC kernels. Thermal design assumes adequate case ventilation to keep the 250 W TDP stable during extended runs. Buyers needing higher memory capacity or newer architecture features should evaluate alternative accelerators.

Highlights

  • Pascal architecture drives 9.3 TFLOPS single-precision compute
  • 16 GB HBM2 memory with 732 GB/s bandwidth
  • 3584 CUDA cores accelerate parallel workloads
  • 7 TFLOPS double-precision for HPC tasks
  • PCI Express 3.0 x16 interface fits dual-slot workstations

Specifications

BrandHP
SeriesHPE ProLiant
ModelQ0E21A
InterfacePCI Express 3.0 x16
Chipset ManufacturerNVIDIA
GPUTesla P100
Core Clock1190 MHz
CUDA Cores3584
Memory Clock715 MHz
Memory Size16GB
Memory Interface4096-bit
Memory TypeHBM2
DirectXDirectX 12.1
OpenGLOpenGL 4.6
System RequirementsMax Power Consumption 250W
Dimensions (L x H)Length: 267 mm(10.5 inches)
Slot WidthDual Slot
Package ContentsOEM Version, No box

Questions about this item

What limits sustained compute on the 3584 CUDA cores first?

The 250 W power ceiling caps clocks before temperature does.

How much heat and fan noise appear when the card runs at 250 W all day?

Expect continuous dual-slot airflow at full 250 W; acoustics depend on chassis fan curves.

Which host connectors and power cables must already be present?

A PCIe 3.0 x16 slot plus one 8-pin and one 6-pin auxiliary power connector.

Also in this aisle

People compared these

Same shelf, same checkout — eight coins and a 30-minute rate lock.