T1#research
SEQUEL and System R — Where SQL Began

Metadata
- Date
- Decade
- 1970s
- Tier
- T1
- Timelines
- A History of Databases
- Sources
- 07
- Connections
- 02
- Tags
- #research
In May 1974, at an ACM SIGFIDET workshop in Ann Arbor, Michigan, two young researchers from IBM's San Jose laboratory presented a sixteen-page paper titled "SEQUEL: A Structured English Query Language." Donald Chamberlin and Raymond Boyce were proposing a way to ask questions of a database that read almost like English. Three years later the language would lose its vowels and become SQL.
What follows is the part of the story that sits between Codd's 1970 relational model paper and Oracle's first commercial SQL product in 1979: who designed the language, what the 1974 paper actually claimed, and what the System R project did with it.
From Codd's Symposium to SQUARE
Chamberlin and Boyce first met Ted Codd in 1972, at a symposium Codd organized at IBM's T.J. Watson Research Center in Yorktown Heights, New York. Both had recently finished PhDs — Chamberlin at Stanford, Boyce at Purdue — and both had been studying the CODASYL Data Base Task Group language, in which a query was a program that navigated a network of pointers.
Chamberlin later described Codd's symposium as a revelation: a query that needed a complex DBTG program could be written in a few lines of one of Codd's relational languages. The two began a game of inventing queries and challenging each other to express them. One that came out of the game was "find names of employees who earn more than their managers."
They also saw a problem. Codd's relational algebra and relational calculus were compact, but they used mathematical notation that could not be typed at a keyboard and concepts — quantifiers, bound variables — drawn from symbolic logic. Their first attempt at something more accessible, called SQUARE, replaced logic with "mappings" over tables. But SQUARE still used subscripts, and so it still could not be typed.
The 1974 SEQUEL Paper
In 1973 Chamberlin and Boyce moved to San Jose Research to join the System R project, and began a new language they called SEQUEL.
The abstract of the 1974 paper states the claim plainly: without resorting to bound variables and quantifiers, SEQUEL identifies a set of simple operations on tabular structures that can be shown to be of equivalent power to the first-order predicate calculus. The intended user was not the programmer alone. The paper names accountants, engineers, architects and urban planners — people unwilling to learn a mathematical notation but willing to learn a structured one.
The paper's first example is the one every SQL user still writes:
SELECT NAME
FROM EMP
WHERE DEPT = 'TOY'
The paper calls the SELECT-FROM-WHERE block the basic component of the language, and suggests that an interactive system might present it to the user as a template with blanks to fill in. It also supports the set functions SUM, COUNT, AVG, MAX and MIN, GROUP BY, and the set operators union, intersection and difference, and closes with an appendix giving the full grammar.
From the beginning SEQUEL was meant to cover data definition as well as queries. A companion report, "Using a Structured English Query Language as a Data Definition Facility," was issued as IBM Research Report RJ1318 in December 1973, but Chamberlin notes that it was never published outside IBM; the query paper became the famous one.
The Death of Ray Boyce
About a month after the Ann Arbor talk, Ray Boyce collapsed at lunch and died of a ruptured brain aneurysm. He was 26. Chamberlin's remembrance at the 1995 System R reunion places his death on Father's Day 1974; Boyce left a wife and an infant daughter. In the short time he had, Boyce also worked with Codd on what is now taught as Boyce–Codd Normal Form.
System R: Three Phases
The 1981 Communications of the ACM paper by the System R team divides the project into three phases, and the dates matter because they are often compressed into a single "1974."
| Phase | When | What happened |
|---|---|---|
| Phase Zero | 1974 – most of 1975 | A single-user interpreter for a subset of SEQUEL, written in PL/I on top of Raymond Lorie's XRM access method. It supported subqueries but not joins. Its code was deliberately thrown away. |
| Phase One | most of 1976 – 1977 | A full multiuser system: the RSS storage layer with locking and logging, and the RDS layer with authorization and access-path selection. |
| Phase Two | 1978 – 1979 | Evaluation at San Jose, at internal IBM sites, and at three customer sites. The first installations took place in June 1977. |
Two engineering decisions from Phase One shaped every later SQL system. Raymond Lorie showed that SQL statements could be compiled into compact System/370 machine-code routines assembled from a library of fragments, which made SQL usable for transaction processing rather than only ad-hoc queries. And the optimizer, in the work IBM credits to Patricia Selinger, chose access paths by estimated cost — the team settled on a weighted sum of CPU time and I/O count.
The language was tested on people as well as machines. Phyllis Reisner, a linguist on the staff, spent several months teaching SEQUEL to students recruited from San Jose State to see whether they could learn it; a human-factors comparison of SQUARE and SEQUEL, co-written with Boyce and Chamberlin, was published at the 1975 National Computer Conference. Chamberlin's later verdict was wry: if you worked hard enough, you could teach SEQUEL to college students, and most of their mistakes had nothing to do with syntax. A revised design, "SEQUEL 2," appeared in the IBM Journal of Research and Development in November 1976, partly informed by the early users.
The customer sites, as the reunion participants remembered them, were Pratt & Whitney Aircraft in Hartford, which used System R for inventory of jet-engine parts; Upjohn in Kalamazoo, which stored the results of clinical experiments for FDA submissions; and later Boeing. At every site the system was installed for study purposes only, not as a supported product.
How SEQUEL Became SQL
In 1977, Chamberlin writes, "because of a trademark issue, the name Sequel was shortened to SQL." At the 1995 reunion he recalled the source of the challenge, with an explicit hedge, as the British company Hawker Siddeley, which said SEQUEL was its registered trademark. The fix was mechanical. In his words: "I think I was the one who condensed all the vowels out of SEQUEL to turn it into SQL," following the pattern of three-letter language names ending in L, such as APL.
The 1981 paper accordingly refers to "the high-level SQL (formerly SEQUEL) language." Neither account gives a date more precise than the year.
From Research Prototype to Industry Standard
The 1981 paper records that the System R research prototype later evolved into SQL/Data System, an IBM product for the DOS/VSE operating system; DB2 on MVS followed. The slow route from laboratory to IBM product, and the startup that shipped SQL first, are covered in the Oracle article. Twelve years after the Ann Arbor paper the language had also become a formal national standard — see the first SQL standard of 1986.
Chamberlin's own assessment, written in 2012, is that SQL was more successful than he and Boyce "had any reason to expect in 1974." It also succeeded in a direction they did not plan. They designed it for ad-hoc questions from planners and other professionals, wanting it simple enough that ordinary people could "walk up and use it." Instead, he observes, it came to be used mostly by trained database specialists to implement repetitive transactions — bank deposits, card purchases, online auctions.
That gap is the clearest measure of what the 1974 paper did. It did not make databases usable by accountants without training. It gave the relational model a syntax that compiled, optimized and ran fast enough to carry the world's transactions — which turned out to matter more.
Questions this page answers
- When was SQL invented?
- Its precursor, SEQUEL, was published in May 1974 in a paper by IBM's Donald Chamberlin and Raymond Boyce. Design work began in 1973, and the name became SQL in 1977.
- Why was SEQUEL renamed SQL?
- Because of a trademark issue in 1977. In Chamberlin's recollection, the British aircraft maker Hawker Siddeley said SEQUEL was its registered trademark, so the vowels were dropped to make SQL.
- Was System R a product?
- No. It was a research system at IBM's San Jose lab, and every installation from June 1977 on was experimental only. It later evolved into SQL/Data System, an IBM product for DOS/VSE.
Sources
The paper itself (Wayback capture of the PDF IBM Almaden once hosted): the abstract's claims, the SELECT-FROM-WHERE block, intended users, relation to SQUARE. ACM DL metadata mislabels the workshop as 1976; it was held in 1974
Dates of Phases Zero, One and Two; the PL/I interpreter on XRM; the prototype without joins; RSS and RDS; first installations in June 1977; evolution into SQL/Data System
Codd's 1972 symposium, SQUARE, the 1973 move to San Jose, Boyce's death at 26, the 1977 renaming, 'walk up and use it'
Participants' transcript: the Hawker Siddeley trademark, dropping the vowels, Reisner's user experiments, the Pratt & Whitney, Upjohn and Boeing joint studies, the date of Boyce's death
SecondaryThe relational database — IBM Heritage
System R begun in 1973; Selinger's cost-based optimizer; Lorie's compiler
TertiaryIBM System R — Wikipedia
Last updated: