"Big data" started as a buzzword about size. Fifteen years later its real legacy is architectural: the way organisations store, process and use data has been rebuilt at least three times. This article traces that transformation, from the enterprise data warehouse through Hadoop data lakes and cloud warehouses to today's lakehouses and real-time, AI-ready platforms, and explains what each shift solved and what it broke.
What big data meant in the first place
The usual definition is the "Vs": volume (more data than one server can handle), velocity (data arriving continuously), and variety (logs, text, images and sensor readings, not just rows and columns). Later lists added veracity (how trustworthy it is) and value (whether it is worth keeping). The Vs describe the problem. The interesting story is how platforms changed to handle it. For the management side of the same story, see my summary of McAfee and Brynjolfsson's "Big Data: The Management Revolution".
Era 1: the enterprise data warehouse
From the 1990s, analytics meant a central relational warehouse, designed with Inmon's normalised approach or Kimball's star schemas. ETL tools extracted data from operational systems every night, applied business rules and loaded clean facts and dimensions; reporting tools queried them. Large organisations ran warehouses on dedicated appliances where storage and compute were bought together.
What it solved: one consistent, historical version of the business. Where it struggled: expensive scaling, rigid schemas that took months to change, and no good home for semi-structured or very large data such as web logs.
Era 2: Hadoop and the data lake
Hadoop, released as an Apache project in 2006, brought distributed storage (HDFS) and processing (MapReduce) on clusters of commodity servers. Storing everything suddenly became cheap. The idea of a data lake, credited to James Dixon when he was CTO of Pentaho, was to keep raw data in its native form and decide how to use it later. Tools such as Hive added SQL on top, and many companies ran hybrid designs with Hadoop feeding the warehouse, like the one I described in 2015 in Data Warehouse Architecture: Traditional ETL vs Big Data Hybrid.
What it solved: cost and scale for raw and semi-structured data. Where it struggled: clusters were hard to operate, MapReduce was slow for interactive queries, and without governance many lakes turned into "data swamps" nobody trusted.
Era 3: cloud data warehouses
Cloud warehouses such as Amazon Redshift, Google BigQuery and Snowflake changed the economics again. Storage and compute were separated: data lives once in cheap cloud storage, and compute is rented by the second or by the query. SQL came back as the main language, and loading raw data first and transforming it inside the warehouse (ELT) became the norm; see ETL vs ELT. For how one of these platforms works internally, see Snowflake architecture explained.
What it solved: elastic scale with little administration, fast SQL, and analysts who could own transformations. Where it struggled: data locked in proprietary formats, machine-learning workloads still copied data out to separate lakes, and cost could grow quietly with usage.
Era 4: the lakehouse
The lakehouse combines the lake's open, cheap storage with the warehouse's reliability. The enabling technology is open table formats: Delta Lake, Apache Iceberg and Apache Hudi. They add a transaction log and table metadata on top of Parquet files in object storage, which brings ACID transactions, schema enforcement, time travel and efficient updates to plain files. Several engines (Spark, Trino, warehouse engines) can read the same tables, so BI and machine learning work from one copy of the data.
What it solves: one platform for SQL analytics and data science, open formats that reduce lock-in. What to watch: more moving parts (catalogs, file compaction, format versions) and governance that must span several engines.
Era 5: real-time and AI-ready platforms
Three developments define the current stage:
- Streaming as a normal input. Event logs such as Apache Kafka and change data capture feed platforms continuously, so dashboards and models can use data that is seconds old where the business needs it.
- Semantic layers and governance. With many tools consuming the same data, metrics are defined once and access is controlled centrally; see What is a semantic layer?.
- AI on governed data. Predictive models and, now, large language models run on curated tables, from risk scores (like my readmission prediction tutorial) to assistants that answer questions in plain English; see Building an AI analytics assistant.
The five eras compared
| Era | Storage | Processing | Main users | Main limitation |
|---|---|---|---|---|
| Enterprise DW | Relational, on-premise | ETL tools, SQL | BI developers, analysts | Cost and rigidity |
| Hadoop data lake | HDFS on commodity servers | MapReduce, Hive, later Spark | Engineers | Operational complexity, swamps |
| Cloud warehouse | Managed columnar storage | Elastic SQL engines | Analysts, analytics engineers | Proprietary formats, cost control |
| Lakehouse | Open table formats on object storage | Spark, SQL engines | Engineers, analysts, data scientists | More components to govern |
| Real-time / AI-ready | Lakehouse or warehouse plus streams | Streaming + batch + ML + LLMs | Everyone, including AI tools | Governance and trust at scale |
What this means for your organisation
- You probably don't need the newest era everywhere. A cloud warehouse with ELT and good tests covers most reporting needs. Add streaming or a lakehouse when a concrete requirement appears.
- Modelling skills transfer. Star schemas, conformed dimensions and data quality rules from the warehouse era are as relevant on a lakehouse as they were on an appliance.
- Governance is now the bottleneck. Storage and compute are cheap; knowing which data is correct, who may use it and what each metric means is the hard part.
- AI raises the stakes. A model or assistant built on inconsistent data produces confident wrong answers faster.
Summary
Big data transformed data platforms in stages: the enterprise warehouse gave consistency, Hadoop gave cheap scale, cloud warehouses gave elasticity and brought SQL back, lakehouses unified analytics and data science on open formats, and today's platforms add real-time data, semantic layers and AI. For how these pieces fit together in a current design, read the modern data engineering architecture guide; for the analytics techniques that run on top, see Big Data Analytics explained.
The Big Data Revolution has transformed the way organizations collect, store, process, and analyze data. With the rapid growth of the internet, social media, mobile devices, IoT sensors, and cloud computing, businesses now generate massive amounts of structured, semi-structured, and unstructured data. Traditional database systems are often unable to handle the volume, velocity, and variety of this data, leading to the adoption of big data technologies such as Hadoop, Spark, and cloud-based data platforms. These technologies enable organizations to process large datasets efficiently and gain valuable insights in real time.
ReplyDeleteThe article provides a useful introduction to Big Data and explains why organizations increasingly depend on large volumes of structured and unstructured information. Big data analytics helps businesses process complex datasets, identify trends and correlations, and extract actionable insights that can support strategic decisions and improve operational efficiency.
ReplyDelete