T1#market#regulation

NVIDIA H100 — The Backbone of AI Compute

Four NVIDIA H100 PCIe cards lined up
Source极客湾Geekerwan (Wikimedia Commons) · CC BY 3.0 · View on Commons ↗

Metadata

Date
Decade
2020s
Tier
T1
Sources
14
Connections
06
Tags
#market#regulation

The NVIDIA H100 is a data-centre GPU that NVIDIA announced at GTC on 22 March 2022. It was the first product built on the new Hopper architecture, named for the US computer scientist Grace Hopper. At the announcement NVIDIA said it would be available starting in the third quarter; on 20 September 2022 it announced that the H100 was in full production, with partners planning to roll out the first wave of products in October, and with H100 instances on AWS, Google Cloud, Microsoft Azure and Oracle Cloud starting the following year. Anyone asking "when" should hold on to that gap of roughly six months between announcement and shipment.

The H100 is not a graphics card in the consumer sense. It was designed as the compute part for training and serving large language models. When ChatGPT appeared that November and demand for generative-AI computing surged, the H100 became one of the parts that absorbed it.

H100 specifications

These are the launch-time figures from NVIDIA's announcement and its architecture blog.

H100 SXM5H100 PCIe
Transistors80 billion (GH100 die, TSMC 4N)Same
Die size814 mm²Same
Enabled SMs132114
Memory80 GB HBM380 GB HBM2e
Memory bandwidth3 TB/s2 TB/s
TDP700 W350 W
GPU interconnectFourth-generation NVLink, 900 GB/s totalPCIe Gen5

The full GH100 die has 144 SMs and 60 MB of L2 cache; the shipping H100 disables part of it and carries 50 MB of L2. NVIDIA's current product page lists the H100 SXM at 3.35 TB/s of memory bandwidth and 3,958 TFLOPS of FP8 Tensor Core throughput with sparsity. The H100 NVL added later, with 94 GB at 3.9 TB/s, is a different product from the PCIe card of the launch.

The Transformer Engine and FP8

The headline feature was the Transformer Engine. NVIDIA describes it as a combination of software and Hopper Tensor Core hardware designed to accelerate Transformer training and inference, choosing dynamically, layer by layer, between 8-bit floating point (FP8) and 16-bit arithmetic.

FP8 comes in two formats: E4M3, with four exponent bits and three mantissa bits, and E5M2, with five and two. Halving the bit width lets the same silicon do more operations, at a cost in precision; the Transformer Engine's job is to watch how much precision each layer can afford to lose.

Hopper also added thread block clusters, which let threads cooperate across several SMs; a Tensor Memory Accelerator (TMA) for moving large tensors; and DPX instructions for dynamic-programming algorithms.

How to read "up to 9x" and "up to 30x"

The numbers that dominated the launch were NVIDIA's own comparisons against the previous generation, the Ampere-based A100: training a 395-billion-parameter Mixture of Experts model "up to 9x faster", and chatbot inference on Megatron 530B with "up to 30x higher throughput than the previous generation". The release does not state the precision or system configuration behind either figure.

NVIDIA's architecture blog gives comparisons with clearer terms: three times the A100's FP64 and FP32 processing rate per clock, and FP8 Tensor Core throughput 6.4 times the A100's FP16. The second of those compares two different precisions. The same pattern recurred with the successor, Blackwell B200.

From announcement to shipment

DateEvent
22 March 2022Hopper and H100 announced at GTC; availability promised from the third quarter
26 August 2022US government imposes a licence requirement on A100 and H100 exports to China (including Hong Kong) and Russia
20 September 2022Full production announced; first partner products planned for October, cloud instances from the next year
17 October 2023US rules expanded to cover the H800, A800 and others (made effective immediately on 23 October)
13 November 2023H200 announced with 141 GB of HBM3e; shipping from the second quarter of 2024

At the production announcement NVIDIA expected more than 50 server models from major makers by the end of the year, opened orders for the DGX H100 with eight H100s, and bundled a five-year NVIDIA AI Enterprise licence with H100 for mainstream servers. Early adopters it named included Los Alamos National Lab, the Swiss National Supercomputing Centre (CSCS) and the University of Tsukuba.

Export controls

The H100 was a geopolitical object before it shipped. According to NVIDIA's August 2022 filing with the SEC, the US government told the company on 26 August that a new licence requirement, effective immediately, applied to any future export of the A100 and the "forthcoming H100" to China (including Hong Kong) and Russia. The stated reason was the risk of diversion to military end use, and NVIDIA said about $400 million of expected third-quarter sales to China might be affected. A follow-up filing dated 1 September said exports needed to continue H100 development had been authorised.

NVIDIA offered cut-down parts such as the H800 for China, but the October 2023 rule change brought those under licence too. The government moved the effective date from 30 days out to immediate on 23 October, and NVIDIA disclosed that, given worldwide demand, it did not expect a near-term meaningful impact on its results. In January 2025 DeepSeek-R1 was built on a base model trained on those H800s. In December 2025 President Trump was reported to have approved sales of the successor H200 to vetted Chinese customers with a 25% fee, excluding Blackwell and Rubin.

Demand, and what followed

NVIDIA's revenue for fiscal 2024 (the year to January 2024) was $60.9 billion, up 126%, with Data Center revenue up 217% to $47.5 billion. On 18 June 2024 NVIDIA briefly became the world's most valuable company (details).

The product line continued with the H200 (announced 2023) and Blackwell B200 (2024), and on 31 May 2026 NVIDIA announced that its successor platform, Vera Rubin, was ramping into full production. Underneath all of them is the parallel programming model NVIDIA first released as CUDA in 2007.

The period around it is on the hardware timeline and the AI timeline; the company and its chief executive have their own pages, NVIDIA and Jensen Huang.

Questions this page answers

What is the H100, and when was it announced?
It is NVIDIA's data-centre GPU and the first product of the Hopper architecture. It was announced at GTC on 22 March 2022, with availability promised from the third quarter. Its main job is training and serving AI models such as large language models.
When did the H100 ship?
NVIDIA announced full production on 20 September 2022, with partners planning to roll out the first products in October. Cloud instances on AWS, Google Cloud, Microsoft Azure and Oracle Cloud were to start in 2023.
What are the H100's main specifications?
At launch: an 80-billion-transistor die on TSMC 4N; the SXM5 version with 132 SMs, 80 GB of HBM3 at 3 TB/s and a 700 W TDP, and the PCIe card with 114 SMs, 80 GB of HBM2e at 2 TB/s and 350 W. NVIDIA's current product page lists the SXM version at 3.35 TB/s.
Can the H100 be exported to China?
On 26 August 2022 the US government imposed a licence requirement on A100 and H100 exports to China (including Hong Kong) and Russia. The October 2023 rule change extended it to cut-down China parts such as the H800.

Sources

  1. PrimaryNVIDIA Announces Hopper Architecture, the Next Generation of Accelerated Computing — NVIDIA Newsroom, Mar 22, 2022

    The GTC announcement; 80 billion transistors, TSMC 4N, HBM3 and 3 TB/s; NVIDIA's up-to-9x and up-to-30x comparisons; availability from the third quarter

    Accessed 2026-10-05

  2. PrimaryNVIDIA Hopper Architecture In-Depth — NVIDIA Technical Blog, Mar 22, 2022

    SXM5 versus PCIe (SMs, HBM3 vs HBM2e, bandwidth, TDP), die size, the two FP8 formats, the Transformer Engine, and comparisons with the A100

    Accessed 2026-10-05

  3. PrimaryNVIDIA Hopper in Full Production — NVIDIA Newsroom, Sep 20, 2022

    Full production, October partner products, cloud from 2023, DGX H100 orders, and early adopters including the University of Tsukuba

    Accessed 2026-10-05

  4. PrimaryNVIDIA H100 Tensor Core GPU — NVIDIA (product page and specifications)

    Current specifications (SXM at 3.35 TB/s, 3,958 TFLOPS FP8 with sparsity, H100 NVL with 94 GB)

    Accessed 2026-10-05

  5. PrimaryNVIDIA Hopper Architecture — NVIDIA

    Accessed 2026-10-05

  6. PrimaryNVIDIA Corporation Form 8-K, Aug 31, 2022 (licence requirement for A100 and H100 exports to China and Russia) — U.S. Securities and Exchange Commission

    The 26 August 2022 licence requirement, China including Hong Kong and Russia, and the roughly $400 million at stake

    Accessed 2026-10-05

  7. PrimaryNVIDIA Supercharges Hopper, the World's Leading AI Computing Platform — NVIDIA Newsroom, Nov 13, 2023

    The H200 announcement (141 GB of HBM3e at 4.8 TB/s, shipping from Q2 2024)

    Accessed 2026-10-05

  8. TertiaryHopper (microarchitecture) — Wikipedia

    Accessed 2026-10-05

Last updated:

Share