Detailed · 9 events

A History of Databases

From Codd's 1970 relational model paper through the Oracle/MySQL/PostgreSQL golden age of RDBMS, the 2009 NoSQL revolt, Snowflake's cloud data warehouse, and the LLM era's vector databases. How structured-data storage and retrieval shifted from application foundation to AI memory.

1950s

  1. EVT.001T1IBM 305 RAMAC — The First Commercial Hard Disk DriveA History of Semiconductors and HardwareA General History of Information TechnologyRead more

1970s

  1. Aerial view of IBM Research Almaden, successor to the IBM San Jose Research Laboratory where Codd worked
    Dicklyon (Wikimedia Commons) · CC BY-SA 4.0 · Commons ↗

    E. F. Codd of IBM's San Jose Research Lab published 'A Relational Model of Data for Large Shared Data Banks' in CACM 13(6), 377–387. Grounded in set theory and first-order predicate logic, it defined n-ary relations and five operations on them (permutation, projection, join, composition, restriction) and put 'data independence'—applications not depending on how data is physically stored—at the centre of the problem. As a fundamental challenge to the then-dominant hierarchical and network DBMS (IMS, IDS), it became the theoretical origin of every RDBMS that followed, from Oracle and DB2 through SQL Server, MySQL, and PostgreSQL. The terms 'relational algebra' and 'relational calculus' are not in this paper; they come from its 1972 sequel, Relational Completeness of Data Base Sublanguages. Codd received the ACM Turing Award in 1981.

    Related people
    Edgar F. Codd
    Related organizations
    International Business Machines (IBM)

    Read more

  2. Larry Ellison, co-founder of Oracle, speaking at Oracle OpenWorld in 2010
    Ilan Costica (Wikimedia Commons) · CC BY-SA 3.0 · Commons ↗

    Software Development Laboratories—founded by Larry Ellison, Bob Miner, and Ed Oates in 1977, renamed Relational Software, Inc. in 1979, then Oracle Systems and finally Oracle Corporation—shipped Oracle V2 in June 1979, the first commercial SQL-based relational DBMS on the market. The company that turned Codd's 1970 relational model paper into a product was not IBM but this startup, writing in PDP-11 assembler. The name came from a CIA project the three had worked on at Ampex, and the CIA became an early customer. Rebuilding Version 3 in C in 1983 opened the market the relational model had implied: a DBMS not tied to any hardware vendor. Revenue for the year ended 31 May 2026 was US$67.4 billion.

    Related people
    Larry Ellison
    Related organizations
    Oracle Corporation

    Read more

1990s

  1. Oracle Corporation (Wikimedia Commons, vectorised by Vulphere) · Public Domain (text logo, ineligible for copyright); MySQL is a trademark of Oracle · Commons ↗

    Michael 'Monty' Widenius, David Axmark, and Allan Larsson founded MySQL AB in Sweden and put out the first, internal version of the MySQL RDBMS. Its licence was not the GPL but a house one—free to use, but paid for commercial use and for the Windows build; the GPL-plus-commercial dual licence arrived only in 2000. As the 'M' in the LAMP stack (Linux, Apache, MySQL, PHP/Perl/Python), MySQL powered the Web boom of the 2000s. Sun Microsystems acquired MySQL AB for about US$1 billion in 2008; when Oracle closed its purchase of Sun in January 2010, MySQL passed to Oracle. Widenius had already forked MariaDB in 2009 to keep an open successor line alive.

    Related organizations
    Oracle Corporation

    Read more

  2. Daniel Lundin (Wikimedia Commons) · BSD · Commons ↗

    On 9 July 1996 Marc Fournier committed 'Postgres95 1.01 Distribution - Virgin Sources' into a newly created repository—the moment the lineage of the UC Berkeley POSTGRES project (begun 1986 under Michael Stonebraker, wound down at Version 4.2) passed to a development community outside the university. Later that same year the project was renamed PostgreSQL and its version numbering was set to 6.0 to resume the Berkeley sequence; 6.0 itself shipped on 29 January 1997. Built around ACID compliance, MVCC, and extensibility, it spawned PostGIS, TimescaleDB, and pgvector—and in Stack Overflow's Developer Survey it took the most-used database slot from MySQL in 2023 and has since held first place for most-used, most-admired, and most-desired alike.

    Read more

2000s

  1. MongoDB, Inc. (Wikimedia Commons) · Public Domain (ineligible for copyright); MongoDB is a trademark · Commons ↗

    10gen (later MongoDB Inc.), founded by Dwight Merriman, Kevin P. Ryan and Eliot Horowitz, made the first public release of MongoDB—a document-oriented NoSQL database that stored BSON-formatted JSON documents instead of rows and columns, with schemaless modelling and horizontal scalability. Version 1.0 shipped that August. It became the public face of the 'NoSQL movement', went public on NASDAQ in 2017, moved the Community Server from AGPLv3 to the SSPL in 2018, and reported US$2.46 billion of revenue in the fiscal year ended January 2026—establishing the document model as a durable choice in a DBMS market that had been dominated by relational systems.

    Related terms
    IaaS / PaaS / SaaS

    Read more

2010s

  1. AWS announced the limited preview of Amazon Redshift at the first-ever re:Invent (general availability followed in February 2013). Built on MPP and columnar technology licensed from ParAccel, the cloud-native analytical data warehouse was priced, with reserved instances, at under US$1,000 per terabyte per year—which the press release called one tenth the price of most data warehousing solutions then available. By overturning the assumption that a data warehouse meant buying dedicated hardware, Redshift—together with BigQuery, made publicly available in May 2012—established the 'cloud DWH' category, opening the path that Snowflake and Databricks would later disassemble and re-architect.

    Related organizations
    Amazon

    Read more

  2. Snowflake (Wikimedia Commons) · Public domain · Commons ↗

    Snowflake Computing—founded in 2012 by Benoit Dageville, Thierry Cruanes, and Marcin Żukowski, veterans of Oracle and Vectorwise—came out of stealth on 21 October 2014, announcing the Snowflake Elastic Data Warehouse and US$26 million in total funding. By separating storage (object stores such as S3) from compute (virtual warehouses), it built a cloud-native DWH where compute could be spun up only as needed. The product reached general availability on 23 June 2015. Its multi-cluster, shared-data architecture powered explosive growth: the NYSE listing on 16 September 2020 closed day one at $253.93 for a market value of roughly US$70.4 billion—the largest software IPO to that date. In FY2026 (year ended 31 January 2026) revenue was US$4.68 billion across 13,328 customers.

    Read more

2020s

  1. Gknor (Wikimedia Commons) · CC BY-SA 4.0 · Commons ↗

    On 3 May 2023 AWS announced pgvector support in Amazon RDS for PostgreSQL (15.2 and later). Supabase had shipped it at Launch Week 7 in April; Google Cloud SQL and AlloyDB followed on 27 June and Amazon Aurora in July, and the major managed PostgreSQL services lined up behind the extension. pgvector stores embeddings—high-dimensional vector representations of text or images—in ordinary PostgreSQL tables and searches them by L2 distance, inner product, or cosine distance; it already shipped an IVFFlat index in v0.1.0 in April 2021. HNSW, the index the purpose-built vector databases standardised on, arrived in v0.5.0 on 28 August 2023. As RAG demand surged after ChatGPT's November 2022 release, pgvector's 'add an extension to your existing PostgreSQL' approach went head to head with purpose-built vector databases such as Pinecone, Weaviate, and Chroma.

    Related organizations
    Google · Microsoft Corporation · Amazon
    Related terms
    Embedding · Large Language Model (LLM) · Retrieval-Augmented Generation (RAG)

    Read more