T1#research#award
Codd's Relational Model Paper — Giving Databases a Mathematical Foundation

Metadata
- Date
- Decade
- 1970s
- Tier
- T1
- Timelines
- A History of Databases
- Sources
- 09
- Connections
- 02
- Tags
- #research#award
In June 1970, Edgar Frank Codd, a researcher at IBM's San Jose Research Laboratory (today IBM Research Almaden), published "A Relational Model of Data for Large Shared Data Banks" in Communications of the ACM, volume 13, issue 6, pages 377–387. It went on to become one of the most cited papers in computer science.
That month, the new technical genre called "the database" was given a mathematical foundation.
Databases in 1970 — A World of Trees and Pointers
Before the paper, two families of systems handled bulk data. One was the hierarchical DBMS, exemplified by IBM's IMS, which stored records as a tree of parent-child relationships. The other was the network DBMS descended from Charles Bachman's IDS and codified by CODASYL, in which records were joined by explicit pointers into a graph.
Both shared a problem. To write a query, the application had to know how the data was physically stored—which tree node, which pointer to follow. Change the storage layout and you rewrote the application. The paper's abstract opens on exactly this point: "Future users of large data banks must be protected from having to know how the data is organized in the machine (the internal representation)."
The Paper's Claim — Sets and Logic Are Enough
Codd's paper runs to eleven pages. Its argument collapses into two claims.
First claim: all data can be represented as 'relations'—mathematical sets. A relation is defined by its attributes (columns), with rows (tuples) belonging to that relation. Not files, not trees, not pointer webs, but collections of sets. That is the abstraction proposed.
Second claim: operations on relations can be expressed with set operations and first-order predicate logic. The 1970 paper defines five: permutation, projection, join, composition, and restriction. On that basis Codd argued that if the relations are in normal form, a first-order predicate calculus suffices, and proposed "a universal data sublanguage based on an applied predicate calculus."
This is easy to get wrong. The terms "relational algebra" and "relational calculus" do not appear in the 1970 paper. Naming the two, and showing that any expression of the calculus can be mechanically reduced to one in the algebra—relational completeness—is the work of a second paper two years later, "Relational Completeness of Data Base Sublanguages" (Courant Computer Science Symposia 6, 1972). SQL's SELECT-FROM-WHERE is closer to the calculus side of that pair.
The natural consequence is data independence. Section 1.2 credits the data description tables of contemporary systems as "a major advance toward the goal of data independence," then enumerates what still binds the application: ordering dependence, indexing dependence, access path dependence. Let the application state only what data it wants, in the language of sets and logic, and leave the DBMS to compile that into physical access paths. Change the storage layout and the queries do not move.
Foreshadowing ACID
Codd's paper does not use the term ACID (Atomicity, Consistency, Isolation, Durability). What it does supply is primary and foreign keys: a way to declare cross-references between records in the user's own terms. (Functional dependency, the core of normalisation theory, is not in the 1970 paper; Codd develops it from 1971 onward.)
Integrity guarantees were stacked on top of that at System R. Jim Gray, who worked on its transaction processing, named atomicity, consistency, and durability in his 1981 paper "The Transaction Concept: Virtues and Limitations"; isolation was added and the acronym ACID coined by Theo Härder and Andreas Reuter in 1983.
Convinced by Codd's paper, an IBM research group at San Jose began building System R in 1974. From it emerged the SEQUEL language (later renamed SQL), two-phase locking, the recovery log, the cost-based query optimiser—broadly, the building blocks of every modern RDBMS trace their lineage there. It was not the only implementation, though: Michael Stonebraker and Eugene Wong read the System R papers and ran INGRES at Berkeley in parallel from 1973, and its query language QUEL is generally judged to have been truer to Codd's algebra than SQL was.
Why IBM Was Slow
There is an irony the history books record. IBM invented the relational model, but IBM was not the first to commercialise it.
IBM's main business protected IMS, the hierarchical DBMS, and senior management was reluctant to ship a competing product line. The lead in the commercial market therefore went to a company founded in 1977 as Software Development Laboratories, renamed Relational Software, Inc. in 1979 and Oracle Systems Corporation shortly afterwards (Oracle's own timeline says 1982, most secondary accounts 1983): it shipped Oracle V2 in that same year, 1979. IBM's first commercial relational product, SQL/DS, arrived in 1981 for DOS/VSE and VM/CMS; DB2 appeared on MVS in 1983 and became generally available in 1985.
Codd himself spent the 1970s evangelising the relational model inside IBM. In October 1985 he published the famous "Codd's 12 Rules" in Computerworld—thirteen of them, in fact, numbered zero to twelve—openly accusing many products that called themselves "relational" of not being so in any rigorous sense.
The Turing Award and the Legacy
In 1981 Codd received the ACM Turing Award, cited "for his fundamental and continuing contributions to the theory and practice of database management systems."
More than half a century later, the dominant systems in the world database market are all from the relational lineage—Oracle, Microsoft SQL Server, IBM Db2, MySQL, PostgreSQL, and on through the cloud-warehouse era of Snowflake and BigQuery. SQL, a language descended from the relational calculus, remains the lingua franca. Even the "NoSQL" movement of the 2000s was framed as a question of how to deal with relational concepts, not how to escape them.
Codd's eleven pages are a rare case in software engineering of mathematical rigour shaping an entire industry. By defining not how data is stored but what data is, he set a frame the field has worked inside for fifty years.
Questions this page answers
- What is the title of Codd's 1970 paper?
- 'A Relational Model of Data for Large Shared Data Banks', published in Communications of the ACM 13(6), pages 377-387, June 1970.
- Did the 1970 paper introduce relational algebra?
- No. The terms relational algebra and relational calculus come from the 1972 sequel, Relational Completeness of Data Base Sublanguages. The 1970 paper defines n-ary relations and five operations on them.
- Did Codd win the Turing Award?
- Yes, the ACM A.M. Turing Award in 1981.
Sources
TertiaryEdgar F. Codd — Wikipedia
TertiaryIBM System R — Wikipedia
TertiaryIngres (database) — Wikipedia
TertiaryACID — Wikipedia
TertiaryRelational model — Wikipedia
Last updated: