T1

NVIDIA Blackwell B200 — Flagship GPU for the Generative-AI Era

Jensen Huang, NVIDIA CEO who led the Blackwell launch
SourceDaniel Torok / The White House (Wikimedia Commons) · Public domain (US Federal Government work) · View on Commons

Metadata

Date
Decade
2020s
Tier
T1
Sources
06
Connections
04

On 18 March 2024, in San Jose, California, NVIDIA CEO Jensen Huang used the GTC 2024 keynote to unveil the next-generation GPU architecture Blackwell—and its first product, the B200. This was an announcement, not a shipment: NVIDIA said partner products would be available "starting later this year," and volume production would in fact slip further still. The name honours David Harold Blackwell, a mathematician who specialised in game theory and statistics and was the first Black scholar inducted into the National Academy of Sciences.

The launch landed sixteen months after ChatGPT's debut, in the middle of a generative-AI boom that was rewriting the data-centre market. Huang's framing: "Generative AI is the defining technology of our time. Blackwell is the engine to power this new industrial revolution."

B200 Specifications

The B200 packs 208 billion transistors. It is a chiplet design: two reticle-limit dies fabricated on TSMC's custom 4NP process, joined by a 10 TB/s chip-to-chip link and presented as a single logical GPU.

The headline numbers are all NVIDIA's own, and they only mean what their conditions say they mean:

  • FP4: 20 PFLOPS—but that is with 2:4 structured sparsity. Dense is 10 PFLOPS.
  • FP8: 10 PFLOPS sparse (5 PFLOPS dense). Against H100's roughly 4 PFLOPS of sparse FP8 that is about 2.5×—a like-for-like comparison at the same precision.
  • Memory: 192 GB of HBM3e at 8 TB/s as specified at announcement. Shipping parts do not match it: HGX B200 carries 180 GB per GPU, and the GB200 Superchip 372 GB across two GPUs, i.e. 186 GB each.
  • TDP: 1000 W in the HGX B200 configuration, higher in GB200. Either way liquid or very dense air cooling is assumed.

The "5× H100 in FP4" line that circulated at the time should be resisted. H100 has no FP4. That figure sets Blackwell's FP4 against Hopper's FP8—a comparison across precisions, not a statement that Blackwell is five times faster.

The new FP4 (4-bit floating-point) format was introduced to extract maximum throughput from low-precision inference workloads—aimed squarely at collapsing the per-token cost of LLM serving.

GB200 NVL72 — The Rack as a Computer

Blackwell's true protagonist is not the single GPU but the rack-scale system GB200 NVL72.

A single cabinet pairs 36 Grace CPUs with 72 Blackwell GPUs—36 GB200 Grace Blackwell Superchips—fused by fifth-generation NVLink (1.8 TB/s bidirectional per GPU) into one huge shared-memory domain.

NVIDIA described it at launch as acting "as a single GPU with 1.4 exaflops of AI performance and 30TB of fast memory." Unpacked against the current official spec sheet: the 1.4 exaflops is sparse NVFP4 (1,440 PFLOPS; 720 PFLOPS dense), and the 30 TB is 13.4 TB of HBM3e plus 17 TB of Grace-side LPDDR5X. Rack price was estimated at roughly US$2–3 million.

The H100 comparisons are likewise NVIDIA's claims with NVIDIA's conditions attached. The stated figure is "up to a 30x performance increase compared to the same number of NVIDIA H100 Tensor Core GPUs for LLM inference workloads," with cost and energy "up to 25x" lower—the footnote specifies 50 ms time-to-next-token, 5 s time-to-first-token, 32,768 input / 1,024 output tokens, against HGX H100 scaled over InfiniBand. The "4x training" claim is a 1.8-trillion-parameter MoE run: 4,096 HGX H100s versus 456 GB200 NVL72 racks.

The product symbolised a shift from "buying GPUs" to "buying an AI factory by the rack".

Delay — Mask Defect and Thermals

Behind the headline event, Blackwell hit a serious manufacturing wall.

In August 2024, The Information reported that Blackwell would slip at least a quarter on "design flaws". NVIDIA denied a logic-design error; CFO Colette Kress said only that the company "executed a change to the Blackwell GPU mask to improve production yields." Industry analysis attributed the problem to a mismatch in thermal-expansion coefficients between the GPU dies, LSI bridges, RDL interposer and substrate inside the CoWoS-L package, causing warpage, and requiring a re-spin of the top metal layers and bump structure. The planned ramp moved out by several months.

GB200 NVL72 racks then ran into overheating issues; some customers had to redesign their liquid-cooling loops.

Mass production finally began in December 2024, with large-scale shipment ramping through Q1 2025. Even then demand outstripped supply: lead times on new orders ran 6–12 months.

Customers and TAM

The launch release named Amazon Web Services, Dell Technologies, Google, Meta, Microsoft, OpenAI, Oracle, Tesla and xAI as organisations "expected to adopt Blackwell"—a list of prospects, not of orders. On the supply side, AWS, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure were named among the first cloud providers to offer Blackwell instances, alongside CoreWeave, Crusoe, IBM Cloud, Lambda and others.

The demand behind the launch is better measured elsewhere: Microsoft, Alphabet, Amazon and Meta together spent over US$200 billion on capital expenditure in 2024, a large share of it on AI data centres. At GTC 2025 Huang presented an analyst-derived chart putting annual data-centre capex on a path to US$1 trillion by 2028.

Competition — AMD MI300X and Google TPU v5p

NVIDIA's runaway was not entirely uncontested.

AMD MI300X: 153 billion transistors on the CDNA 3 architecture, 192 GB HBM3. Microsoft and Meta announced deployments. The CUDA-ecosystem moat, however, kept AMD perpetually behind on software readiness.

Google TPU v5p: powered Google's own frontier training. Not sold externally—available only via Google Cloud.

AWS Trainium2, Microsoft Maia 100: every cloud provider accelerated its in-house silicon programme. The strategic question—"how long do we keep paying NVIDIA?"—drove the captive-chip push.

Even so, market-research estimates of NVIDIA's share of data-centre AI accelerators through 2024–2025 clustered between 80% and above 90%—the estimates vary, the conclusion does not. Rivals still sat at "an alternative that complements NVIDIA", not "a replacement".

Bubble Concerns

Beneath the euphoria, sceptical voices were building. Could hyperscaler AI capex ever be recovered? Would falling inference costs and the rise of Chinese open-weights models (e.g. DeepSeek) eventually shrink demand for NVIDIA's flagship GPUs? Those questions detonated as actual share-price destruction in January 2025 with the DeepSeek-R1 shock.

On the demand side at least, the sceptics were wrong. Blackwell took the majority of NVIDIA's high-end GPU shipments in 2025; the mid-life GB300 NVL72 (Blackwell Ultra) began shipping in the second half of that year, with the Rubin generation following in 2026—a cadence closer to annual than biennial.

Blackwell B200 thus stands as both the apex product of the generative-AI compute build-out and a symbol of NVIDIA's peak one-firm dominance. From this summit, the story moves on—into an era in which the economics of AI compute themselves are put under question.

Sources

  1. TertiaryBlackwell (microarchitecture) — Wikipedia

    Accessed 2026-08-03

Last updated:

Share