Business Professionals
Power BI | Power Pivot | Power Query | DAX
Cloud Flows | RPA | AI Builder | Copilot
60+ Formulas | Data Stories | Advanced Reporting & Modeling
VB Programming | Report Automation |
MS-Office Automation
Techno-Business Professionals
Power BI | Power Query | Advanced DAX | SQL - Query &
Programming
Microsoft Fabric | Power BI | Power Query | Advanced DAX |
SQL - Query & Programming
Power BI | Power Apps | Power Automate | Copilot Studio | Power Pages | Dataverse
Microsoft Power Apps | Microsoft Power Automate
Power BI | Adv. DAX | SQL (Query & Programming) |
VBA | Python | Web Scrapping | API Integration
Power BI | Power Apps | Power Automate |
SQL (Query & Programming)
Power BI | Adv. DAX | Power Apps | Power Automate |
SQL (Query & Programming) | VBA | Python | Web Scrapping | API Integration
Power Apps | Power Automate | SQL | VBA | Python |
Web Scraping | RPA | API Integration
Technology Professionals
Power BI | DAX | SQL | ETL with SSIS | SSAS | VBA | Python
Power BI | SQL | Azure Data Lake | Synapse Analytics |
Data Factory | Databricks | Power Apps | Power Automate |
Azure Analysis Services
Microsoft Fabric | Power BI | SQL | Lakehouse |
Data Factory (Pipelines) | Dataflows Gen2 | KQL | Delta Tables | Power Apps | Power Automate
Power BI | Power Apps | Power Automate | SQL | VBA | Python | API Integration
Power BI | Advanced DAX | Databricks | SQL | Lakehouse Architecture
Business Professionals
Power BI | Power Pivot | Power Query | DAX
Cloud Flows | RPA | AI Builder | Copilot
60+ Formulas | Data Stories | Advanced Reporting & Modeling
VB Programming | Report Automation |
MS-Office Automation
Techno-Business Professionals
Power BI | Power Query | Advanced DAX | SQL - Query &
Programming
Microsoft Fabric | Power BI | Power Query | Advanced DAX |
SQL - Query & Programming
Power BI | Power Apps | Power Automate | Copilot Studio | Power Pages | Dataverse
Microsoft Power Apps | Microsoft Power Automate
Power BI | Adv. DAX | SQL (Query & Programming) |
VBA | Web Scrapping | API Integration
Power BI | Power Apps | Power Automate |
SQL (Query & Programming)
Power BI | Adv. DAX | Power Apps | Power Automate |
SQL (Query & Programming) | VBA | Web Scrapping | API Integration
Power Apps | Power Automate | SQL | VBA |
Web Scraping | RPA | API Integration
Technology Professionals
Power BI | DAX | SQL | ETL with SSIS | SSAS | VBA
Power BI | SQL | Azure Data Lake | Synapse Analytics |
Data Factory | Azure Analysis Services
Microsoft Fabric | Power BI | SQL | Lakehouse |
Data Factory (Pipelines) | Dataflows Gen2 | KQL | Delta Tables
Power BI | Power Apps | Power Automate | SQL | VBA | API Integration
Power BI | Advanced DAX | Databricks | SQL | Lakehouse Architecture
You come in at 08:15, coffee in hand, and the overnight pipeline is red. The dashboard that should show yesterday’s sales is blank, and a stakeholder is already asking why their report is late. If you’ve ever wondered what does a data engineer do when this happens, this is the day-in-the-life version.
We’ll walk through one realistic Azure scenario: a broken nightly pipeline, a silent schema change in a source system, and the conversation when the business wants answers. Along the way we’ll unpack what an Azure data engineer actually does to keep things running.
Our fictional environment is typical for a mid-sized analytics team:
/raw, /curated, /semanticThe key pipeline:
SalesOrders from Azure SQL DB to raw/sales_orders/YYYY/MM/DD/ (parquet)raw/accounts/curated/sales/ partitioned by OrderDateThe business expectation: by 07:00, the “Daily Sales” dashboard shows complete data for the previous day.
The first thing an Azure data engineer does most mornings is look for red.
You open the Monitor view in Azure Data Factory and see:
PL_Nightly_Sales failed at 02:13.DF_Curated_Sales Data Flow activity.Column 'CustomerSegment' not found in input stream 'SalesOrders'.
This already hints at a schema change in the source.
You drill into the pipeline run:
So ingestion worked, transformation didn’t. That matters for how you respond to stakeholders: data is in the lake, but not yet in the curated layer the dashboard uses.
You open the latest file in ADLS (e.g. via Azure Storage Explorer or Synapse serverless):
SELECT TOP 50 *
FROM OPENROWSET(
BULK 'https://<storage-account>.dfs.core.windows.net/datalake/raw/sales_orders/2026/10/05/*.parquet',
FORMAT = 'PARQUET'
) AS [r];
You confirm:
CustomerSegment is indeed missing.This is the first core task of an Azure data engineer: diagnose where the failure occurred and whether data is safely landed anywhere.
Nobody told you the source system changed. But the data flow was relying on a column that no longer exists.
Typical causes:
You query the source schema directly:
SELECT COLUMN_NAME, DATA_TYPE
FROM INFORMATION_SCHEMA.COLUMNS
WHERE TABLE_NAME = 'SalesOrders'
ORDER BY ORDINAL_POSITION;
You see CustomerSegment is gone, and a new column CustomerTier was added.
From an Azure platform perspective, nothing in ADF auto-adjusts for this. The Copy activity will happily copy whatever the source schema is; the Data Flow fails because its projection expects a column that no longer exists.
You have two competing pressures:
You usually have three options:
CustomerTier instead of CustomerSegment.CustomerTier into the semantic concept the report expects.CustomerSegment from the Data Flow output.CustomerSegment from CustomerTier.What an Azure data engineer actually does all day is make this trade-off repeatedly: short-term continuity vs long-term model integrity.
In most real teams, the morning answer is: get the pipeline green with the least semantic damage, then schedule a proper model change.
Let’s say you choose option 1: adapt to CustomerTier.
In the ADF Data Flow designer, you:
DF_Curated_Sales.SalesOrders so it picks up the new schema.CustomerSegment with CustomerTier.You might add a derived column to keep the semantic name stable:
-- Expression in Data Flow derived column (using Data Flow expression language)
iif(CustomerTier == 'Gold', 'Premium',
iif(CustomerTier == 'Silver', 'Standard', 'Other'))
That derived column is named CustomerSegment in the Data Flow output, so downstream tables and Power BI don’t break immediately.
Important detail: in Data Flows, projection refresh is manual. If you don’t refresh, the Data Flow still thinks CustomerSegment exists, and the run will keep failing.
You trigger a manual rerun of the nightly pipeline for yesterday’s date:
pProcessingDate (string or date).This is where design matters. A resilient Azure pipeline typically:
2026-10-05).You ensure the curated sink writes with overwrite semantics for that partition, or uses a delete-and-insert pattern.
After the rerun:
You now check the Power BI side.
Even with the pipeline fixed, the dashboard can still be wrong or stale. In this scenario, the Power BI dataset refresh depends on the curated layer.
In the Power BI Service:
You see:
The key didn't match any rows in the table.
This often occurs when a schema change hits Power Query or the dataflow.
If you’re using Power BI dataflows on top of curated parquet:
Typical issues:
Table.RenameColumns step referencing CustomerSegment throws an error.Table.SelectColumns step expects CustomerSegment but the column name changed.You update the query to align with the new derived column in curated sales. For example:
let
Source = AzureStorage.DataLake("https://<storage-account>.dfs.core.windows.net", [HierarchicalNavigation=true]),
Curated = Source{[Name="datalake"]}[Data]{[Name="curated"]}[Data]{[Name="sales"]}[Data],
Filtered = Table.SelectRows(Curated, each [OrderDate] >= Date.AddDays(Date.From(DateTime.LocalNow()), -1)),
Selected = Table.SelectColumns(Filtered, {"OrderId", "OrderDate", "CustomerId", "CustomerSegment", "Revenue"})
in
Selected
Here, CustomerSegment is still present because you kept the name stable in the Data Flow. If you had changed it to CustomerTier end-to-end, you’d adjust this step accordingly.
You then trigger a dataset refresh manually. Once it succeeds, the dashboard is live again.
The third core part of the Azure data engineer role is translation: explaining technical causes and realistic guarantees to non-technical stakeholders.
You’ll typically cover three points:
SalesOrders source removed a column the pipeline and report relied on.The details matter. For example, you can explain:
This is often where the stakeholder finally understands what an Azure data engineer does beyond “building pipelines”.
Once the fire is out, the rest of the day is the real job: making sure this happens less often and is easier to fix.
You can implement checks such as:
INFORMATION_SCHEMA.COLUMNS or system views, writes schema snapshots into ADLS, and alerts on differences.Example: a simple ADF pipeline that:
metadata/schemas/salesorders/.This doesn’t prevent schema changes, but it moves the failure to a clear, early stage and makes impact analysis easier.
To support reruns:
ProcessingDate.A simple example: in a Copy activity to parquet, use a dynamic path:
@concat('raw/sales_orders/', formatDateTime(pProcessingDate, 'yyyy/MM/dd'), '/')
This makes correcting a failed run for a specific date straightforward.
The scenario showed why layer separation matters:
An Azure data engineer spends a lot of time enforcing this separation so that failures are localized and recoverable.
Some of the work is social, not technical:
None of this is an Azure feature, but it’s central to the Azure data engineer role.
If you strip away the tooling names, the core activities in this scenario are what define the Azure data engineer role:
The broken overnight pipeline, the schema change, and the late dashboard are not edge cases—they’re the normal rhythm of the job.
If you remember only one thing from this day-in-the-life: make yesterday’s data a parameter. When your Azure pipelines, storage paths, and Power BI queries all hinge on an explicit ProcessingDate, recovering from a broken run becomes a controlled, repeatable operation instead of a scramble through ad-hoc fixes.
This article reflects the shift from ad-hoc pipeline building to disciplined, failure-aware data engineering in Azure environments where nightly loads and business-critical dashboards depend on fragile upstream schemas.
Professionals who want to apply these patterns to their own data can explore Excelgoodies' Data Engineering & BI Azure (On Cloud) programme - taught live by instructors, with certification awarded once a real project is running at work.
Insights compiled through ongoing industry research and discussions within the Excelgoodies Analytics Community.
Azure
New
Next Batches Now Live
Power BI
SQL
Power Apps
Power Automate
Microsoft Fabrics
Azure Data Engineering