Skip to main content

Posts

Showing posts from September, 2026

Big Data: The Management Revolution (HBR 2012): Summary and Review

In October 2012 the Harvard Business Review published "Big Data: The Management Revolution" by Andrew McAfee and Erik Brynjolfsson, then both at MIT. It became one of the most cited business articles on big data, and it still shows up in searches today. This post summarises its argument in plain terms, then looks at what held up, what didn't, and what it means for data teams in the age of cloud platforms and AI. Data-driven organisation Leadership set goals, ask better questions Talent management data scientists, engineers Technology store and process new data Decision making data where decisions happen Company culture evidence over the HiPPO

Building an AI Analytics Assistant: Text-to-SQL with Guardrails

"What was revenue by region last quarter?" Large language models can now turn a question like that into SQL, which makes an AI analytics assistant one of the most requested features on data teams' roadmaps. They are also easy to build badly: an assistant that runs whatever SQL the model writes, on raw tables, with a powerful database user, will eventually give a confident wrong answer or touch data it should not. This article shows the architecture that makes text-to-SQL safe enough for real use, with a runnable guardrail example. User question LLM + semantic layer context SQL validator read-only, allow-list Warehouse read-only role Answer + SQL shown to user Blocked → explain

What Is a Semantic Layer? Define Metrics Once, Use Them Everywhere

Two dashboards show different revenue for the same month. Finance says one number, sales says another, and a meeting that should have been about decisions turns into a debate about whose query is right. The data is not wrong; the definitions are inconsistent. A semantic layer fixes this by defining business metrics once, in one place, and making every tool use those definitions. Warehouse tables fct_orders dim_customer dim_date Semantic layer net_revenue = SUM(gross − discount) orders = COUNT(DISTINCT id) dimensions: region, month row-level access rules BI dashboard Spreadsheet / notebook AI assistant

Apache Airflow for Data Engineers: DAGs, Scheduling and Best Practices

Every data platform needs something that runs jobs in the right order, at the right time, and does something sensible when they fail. Apache Airflow is the most widely used open-source tool for that job. It started at Airbnb in 2014, became a top-level Apache project, and reached version 3 in 2025. This article explains the concepts you need, shows a complete daily pipeline, and lists the practices that keep Airflow pipelines reliable. DAG: daily_sales_load (runs every day at 02:00) extract source → object storage load_raw COPY into raw schema transform dbt build check_quality fail run if data is bad Each box is a task. Arrows are dependencies. Failed tasks retry, then alert; past days can be re-run (backfilled).

Snowflake Architecture Explained, vs Databricks and BigQuery

Snowflake is one of the most widely used cloud data warehouses, and its architecture explains most of its strengths and most of its cost surprises. This article walks through the three layers, the features that follow from them, and how Snowflake compares with the two platforms it is most often weighed against: Databricks and Google BigQuery. Cloud services layer authentication · metadata · query optimizer · access control · transactions Query processing: independent virtual warehouses Loading warehouse XS, auto-suspend Transformation warehouse M, runs dbt at night BI warehouse multi-cluster for concurrency Database storage compressed columnar micro-partitions in cloud object storage (shared by all warehouses)

ETL vs ELT: Differences, Examples and When to Use Each

ETL and ELT use the same three steps (extract, transform, load) in a different order. In ETL the data is cleaned and reshaped before it reaches the warehouse. In ELT the raw data is loaded first and transformed inside the warehouse, using the warehouse's own compute. The order sounds like a detail, but it changes where your logic lives, what you can reprocess, what you pay for, and who can maintain the pipeline. ETL Sources Transform in ETL tool / server Load clean data only Warehouse modelled tables ELT Sources Load raw data as-is Warehouse: raw → Transform with SQL → marts compute of the warehouse does the work

Modern Data Engineering Architecture: A Practical Guide

"Modern data engineering" gets used for everything from a single dbt project to a company-wide streaming platform. Underneath the buzzwords, almost every modern analytics stack has the same six layers: data comes from sources , is ingested , lands in storage , gets transformed into trusted tables, is described by a semantic layer , and is finally consumed by dashboards, models and, increasingly, AI assistants. Two concerns cut across all of them: orchestration and governance . This guide is the map. Each layer gets a short explanation, the decisions that matter, and a link to a deeper article on this blog. Governance: catalog, access control, data quality, privacy Sources Apps & databases SaaS APIs Events / IoT Files Ingestion Batch loads CDC Streaming Storage Object storage + warehouse or lakehouse raw → staging → marts Transformation SQL / dbt Spark Data tests Semantic layer Metrics Dimensions Access rules Consumption BI dashboards ML models AI assistant Orchestratio...