T1
CUDA 1.0 Released — General-Purpose GPU Computing

Metadata
- Date
- Decade
- 2000s
- Tier
- T1
- Sources
- 10
- Connections
- 03
CUDA (Compute Unified Device Architecture) is NVIDIA's way of using its GPUs as general-purpose parallel computers without going through a graphics API. NVIDIA announced it in November 2006 alongside the GeForce 8800 (G80), released version 0.8 as the first public beta in February 2007, and shipped 1.0 in June 2007; the cover of the 1.0 Programming Guide is dated 6/23/2007. The computing base of later deep learning and generative AI runs in a straight line from that June release.
Before CUDA: computing by pretending to draw
Using a GPU for computation is older than CUDA. In the early 2000s researchers recast their calculations as graphics work and had the GPU solve them through a graphics API — the practice known as GPGPU.
The CUDA 1.0 Programming Guide lists the obstacles. The GPU "could only be programmed through a graphics API", which imposed a steep learning curve on newcomers and the overhead of an unsuitable API on non-graphics work. Programs could read from anywhere in DRAM but could not write to arbitrary locations. And some applications were bottlenecked by DRAM bandwidth, leaving the GPU's arithmetic under-used.
The academic attack on those limits came from Stanford's Brook. The SIGGRAPH 2004 paper "Brook for GPUs: Stream Computing on Graphics Hardware", by Ian Buck and colleagues, extended C with simple data-parallel constructs so the GPU could be used as a streaming coprocessor. Buck finished his PhD that same year, 2004, and joined NVIDIA. NVIDIA's own blog describes him as the development lead for Brook and the inventor of CUDA.
November 2006: announced with the G80
On 8 November 2006 NVIDIA unveiled CUDA as a new architecture for computing on its GPUs. The company called it the industry's first C-compiler development environment for the GPU and said it could solve complex computing problems up to 100 times faster than traditional approaches — both NVIDIA's own claims. The supported hardware was the GeForce 8800 launched at the same time, with future Quadro products to follow.
What CUDA 1.0 contained
The 1.0 Programming Guide defines the system this way:
CUDA stands for Compute Unified Device Architecture and is a new hardware and software architecture for issuing and managing computations on the GPU as a data-parallel computing device without the need of mapping them to a graphics API.
The software came in layers: a hardware driver, an API and its runtime, and two maths libraries, CUFFT (fast Fourier transforms) and CUBLAS (linear algebra). Programmers wrote an extension of the C language, which the guide presents as the choice that keeps the learning curve to a minimum.
| CUDA 1.0 (2007) | |
|---|---|
| Supported GPUs | GeForce 8 Series, Quadro FX 5600/4600, Tesla |
| Language | A minimal set of extensions to C |
| Bundled libraries | CUFFT, CUBLAS |
| Maximum threads per block | 512 |
| Shared memory per multiprocessor | 16 KB |
| Warp size | 32 threads |
| Multiprocessors (GeForce 8800 GTX, Tesla C870) | 16 |
| Double-precision floating point | Not native (double demoted to float) |
Work is organised in three levels: threads grouped into blocks, and blocks laid out in a grid. Threads in the same block share data through on-chip shared memory and can synchronise; each multiprocessor issues instructions for a warp of 32 threads at a time. With different names and larger numbers, that structure persists in today's GPUs.
The lack of double precision meant early CUDA could not cover all of scientific computing. For work that single precision could handle — image and signal processing, some physics — uses appeared quickly.
The early releases
The version history in the Windows release notes is short.
| When | Version | Notes |
|---|---|---|
| November 2006 | — | CUDA announced with the GeForce 8800 |
| February 2007 | 0.8 | Initial public beta |
| June 2007 | 0.9, 1.0 | 1.0 is the first full release |
| December 2007 | 1.1 | Listed in the Toolkit archive |
| August 2008 | 2.0 | Listed in the Toolkit archive |
NVIDIA's CUDA Toolkit Archive still lists "CUDA Toolkit 1.0 (June 2007)" as its oldest entry. In September 2026 the top of the same list was 13.4.2.
Competing with an open standard
CUDA runs only on NVIDIA GPUs. OpenCL, proposed by Apple, was ratified and released as version 1.0 by the Khronos Group on 9 December 2008, which billed it as the first open, royalty-free standard for parallel programming across processors; NVIDIA took part in its development.
CUDA stayed dominant all the same. As this site's NVIDIA page argues, what makes switching expensive is less the language than the libraries and tools accumulated on top of CUDA.
Meeting deep learning
CUDA was aimed at scientific computing. The turn came five years later. In 2012 a University of Toronto team trained a network on two GeForce GTX 580 cards with convolution code written in CUDA and won the ILSVRC image-recognition contest by a wide margin (AlexNet); the team released that code with the paper.
From then on deep-learning computation ran on GPUs and CUDA as a matter of course, and the H100 of 2022 and the Blackwell B200 of 2024 descend from it. See also GPU and deep learning in the glossary. The company's story is on NVIDIA; the hardware of the period is on the hardware timeline.
Questions this page answers
- When was CUDA released?
- It was announced with the GeForce 8800 in November 2006, and version 0.8 was the first public beta in February 2007. Version 1.0 followed in June 2007; the 1.0 Programming Guide is dated 23 June.
- What is CUDA?
- It is NVIDIA's hardware and software architecture for using its GPUs as data-parallel computers without mapping the work to a graphics API. Programs are written in an extension of C, supported by a driver, a runtime and maths libraries (CUFFT and CUBLAS).
- Who created CUDA?
- NVIDIA credits Ian Buck, who led development of the Brook stream-computing language at Stanford and joined NVIDIA in 2004, as the inventor of CUDA. Brook was presented in a SIGGRAPH 2004 paper.
- Which GPUs did CUDA 1.0 support?
- The GeForce 8 Series, Quadro FX 5600/4600 and Tesla. None of them handled double precision natively, so the double type was demoted to float.
Sources
PrimaryNVIDIA CUDA Compute Unified Device Architecture Programming Guide, Version 1.0 (6/23/2007) — NVIDIA
The definition, the three prior obstacles, supported GPUs, the software stack, threads/blocks/grids, 512 threads, 16 KB, warp of 32, and double precision
PrimaryNVIDIA CUDA Release Notes, Windows, Version 1.0 — NVIDIA
Revision history (0.8 initial public beta in February 2007; 0.9 and 1.0 in June)
PrimaryIan Buck — author page, NVIDIA Blog
Buck joining NVIDIA in 2004, leading Brook, and NVIDIA's description of him as the inventor of CUDA
SecondaryNvidia Announces Cuda GPU Architecture — Game Developer, Nov 8, 2006
The November 2006 announcement and NVIDIA's first-C-compiler and up-to-100x claims
TertiaryCUDA — Wikipedia
Last updated: