T1#market
Amazon Redshift Announced — The Cloud Data Warehouse Era Begins
Metadata
- Date
- Decade
- 2010s
- Tier
- T1
- Timelines
- A History of Databases
- Sources
- 09
- Connections
- 01
- Tags
- #market
On 28 November 2012, at the inaugural AWS re:Invent conference (27-29 November, The Venetian, Las Vegas), Andy Jassy—then AWS Senior Vice President and later Amazon's CEO—unveiled the limited preview of a new service: Amazon Redshift. It was a cloud-native data warehouse (DWH) promising petabyte-scale analytics at, in the press release's words, "one tenth the price of most data warehousing solutions available to customers today".
It was the moment the cloud arrived seriously in data analytics, the territory that on-premises hardware vendors had ruled for decades.
The DWH Industry in 2012 — An Oligopoly of Expensive Hardware
In the early 2010s the DWH market was dominated by a handful of hardware vendors. Teradata was the leader, followed by Netezza (acquired by IBM in 2010), Greenplum (acquired by EMC in 2010), Vertica (acquired by HP in 2011), and Oracle's Exadata.
What they shared technically was MPP (massively parallel processing) and columnar storage. MPP distributed query execution across many nodes. Columnar storage saved data column by column rather than row by row, sharply cutting the bytes an analytical query (GROUP BY, aggregation, range scans) has to read.
But all of this was heavy, on-premises hardware. The Redshift press release opens by saying as much: self-managed on-premises data warehouses require significant time and resource to administer, especially for large datasets, and the financial cost of building, maintaining, and growing them is flat-out expensive. When you needed more analytics capacity, you ordered another rack and waited for it to arrive and be commissioned.
The ParAccel Licence and the Birth of Redshift
AWS built the service under the internal code name "Cookie Monster". As a base, it chose the MPP and columnar technology of ParAccel, a California DWH vendor. ParAccel had forked PostgreSQL and added column storage and a parallel execution engine, in a product called PADB.
The transfer was a licence, not an acquisition: the press release states plainly that "Amazon Redshift includes technology components licensed from ParAccel". The Register reports that Amazon also led ParAccel's Series E round in July 2011; ParAccel itself was bought by Actian in April 2013. The licensed code was integrated into Amazon's S3/EC2/VPC infrastructure and wrapped as a managed service—Redshift. AWS documentation stated for years that "Amazon Redshift is based on PostgreSQL 8.0.2", so existing psql clients, ODBC/JDBC drivers, and BI tools (Tableau, Looker) could connect unchanged.
28 November 2012 — Announcement at re:Invent
re:Invent 2012 was AWS's first ever conference, announced in May with more than 100 sessions on the programme. The Redshift limited preview announced there carried a price tag that shook the industry: on-demand pricing from US$0.85 per hour for a 2 TB warehouse, or an effective $0.228 per hour—under US$1,000 per terabyte per year—with reserved instances. The press release called that less than one tenth the price of comparable technology then available. Nodes came in 2 TB and 16 TB compressed sizes, and a cluster could scale to 100 of them.
The performance figures were the vendor's own. Raju Gulabani, AWS's Vice President of Database Services, said in the release that while actual performance would vary with each customer's queries, AWS's internal tests had shown over 10 times the performance of standard relational data warehouses. Erik Selberg, who managed Amazon's own data warehouse team, added that early estimates put Redshift's cost well under a tenth of their existing solution's.
Plus the cloud-native advantages: spin a cluster up in a few clicks, scale node count up and down later, never buy or install hardware. Demand at preview outran AWS's plans—within about three days the team saw ten times the demand it had budgeted for the whole first year, and scrambled to accelerate hardware orders (Amazon Science interview). General availability came in February 2013.
The Birth of "Cloud DWH" as a Category
More than as an individual product, Redshift mattered because it created the "cloud DWH" category. Google's BigQuery—Dremel's commercial form, serverless—had gone to limited preview in November 2011 and was made publicly available on 1 May 2012, half a year ahead of Redshift. Microsoft's Azure SQL Data Warehouse (later Synapse Analytics) was announced and previewed in 2015 and reached general availability on 12 July 2016.
Three-way cloud competition triggered seismic shifts in the on-premises DWH market. ParAccel, whose technology underpinned Redshift, was sold to Actian in April 2013; IBM's PureData System for Analytics—Netezza—reached end of support in 2019, later returning as Netezza Performance Server on Cloud Pak for Data. Greenplum and Vertica retreated into enterprise niches.
Redshift's Evolution from 2017
Early Redshift had storage and compute coupled together: to add capacity, you added nodes. This was the exact attack surface that Snowflake—out of stealth in October 2014, generally available in June 2015—would target with its storage-compute separation.
AWS responded with Redshift Spectrum (2017), which queries S3 data directly; the RA3 node family (2019) with Managed Storage, effectively separating storage and compute; and Redshift Serverless, previewed in November 2021 and generally available in July 2022. It has absorbed Snowflake's design ideas while leaning on its strength: tight integration with the rest of the AWS ecosystem.
In the mid-2020s the cloud DWH market centres on Redshift, Snowflake, and BigQuery. Teradata pursues VantageCloud to push into the cloud as well, but the centre of gravity has moved off premises.
From Buying a Warehouse to Renting One
The Redshift announcement was the inflection point from buying a DWH to renting one by the hour. An industry that had operated for decades on the model of purchasing dedicated hardware, waiting for delivery, and hiring DBAs to run it shifted to provisioning a cluster in minutes. What Rahul Pathak recalls from the launch is exactly that pairing: disbelief at the price, and the fact that you could provision a data warehouse in minutes instead of months.
Years later, on the same cloud-native foundation, Snowflake would catch up to Redshift with an architecture that fully separated storage and compute. The evolution of the cloud DWH began at the door Redshift opened.
Sources
TertiaryAmazon Redshift — Wikipedia
TertiaryParAccel — Wikipedia
Last updated: