Skip to main content

Posts

Explore by topic

Analytics & BI

Predictive analytics, semantic layers and reporting architecture.

AI for Data

AI analytics assistants, LLMs over enterprise data and RAG for business data.

Big Data: The Management Revolution (HBR 2012): Summary and Review

In October 2012 the Harvard Business Review published "Big Data: The Management Revolution" by Andrew McAfee and Erik Brynjolfsson, then both at MIT. It became one of the most cited business articles on big data, and it still shows up in searches today. This post summarises its argument in plain terms, then looks at what held up, what didn't, and what it means for data teams in the age of cloud platforms and AI. Data-driven organisation Leadership set goals, ask better questions Talent management data scientists, engineers Technology store and process new data Decision making data where decisions happen Company culture evidence over the HiPPO
Recent posts

Building an AI Analytics Assistant: Text-to-SQL with Guardrails

"What was revenue by region last quarter?" Large language models can now turn a question like that into SQL, which makes an AI analytics assistant one of the most requested features on data teams' roadmaps. They are also easy to build badly: an assistant that runs whatever SQL the model writes, on raw tables, with a powerful database user, will eventually give a confident wrong answer or touch data it should not. This article shows the architecture that makes text-to-SQL safe enough for real use, with a runnable guardrail example. User question LLM + semantic layer context SQL validator read-only, allow-list Warehouse read-only role Answer + SQL shown to user Blocked → explain

What Is a Semantic Layer? Define Metrics Once, Use Them Everywhere

Two dashboards show different revenue for the same month. Finance says one number, sales says another, and a meeting that should have been about decisions turns into a debate about whose query is right. The data is not wrong; the definitions are inconsistent. A semantic layer fixes this by defining business metrics once, in one place, and making every tool use those definitions. Warehouse tables fct_orders dim_customer dim_date Semantic layer net_revenue = SUM(gross − discount) orders = COUNT(DISTINCT id) dimensions: region, month row-level access rules BI dashboard Spreadsheet / notebook AI assistant

Apache Airflow for Data Engineers: DAGs, Scheduling and Best Practices

Every data platform needs something that runs jobs in the right order, at the right time, and does something sensible when they fail. Apache Airflow is the most widely used open-source tool for that job. It started at Airbnb in 2014, became a top-level Apache project, and reached version 3 in 2025. This article explains the concepts you need, shows a complete daily pipeline, and lists the practices that keep Airflow pipelines reliable. DAG: daily_sales_load (runs every day at 02:00) extract source → object storage load_raw COPY into raw schema transform dbt build check_quality fail run if data is bad Each box is a task. Arrows are dependencies. Failed tasks retry, then alert; past days can be re-run (backfilled).

Snowflake Architecture Explained, vs Databricks and BigQuery

Snowflake is one of the most widely used cloud data warehouses, and its architecture explains most of its strengths and most of its cost surprises. This article walks through the three layers, the features that follow from them, and how Snowflake compares with the two platforms it is most often weighed against: Databricks and Google BigQuery. Cloud services layer authentication · metadata · query optimizer · access control · transactions Query processing: independent virtual warehouses Loading warehouse XS, auto-suspend Transformation warehouse M, runs dbt at night BI warehouse multi-cluster for concurrency Database storage compressed columnar micro-partitions in cloud object storage (shared by all warehouses)