Excelgoodies logo +31 97 010285556

LEARN THIS HANDS ON

Full Stack BI (On-Cloud)

. Live Online FILLING FAST
View all upcoming batches
What Does an Azure Data Engineer Actually Do All Day?

What Does an Azure Data Engineer Actually Do All Day?

You come in at 08:15, coffee in hand, and the overnight pipeline is red. The dashboard that should show yesterday’s sales is blank, and a stakeholder is already asking why their report is late. If you’ve ever wondered what does a data engineer do when this happens, this is the day-in-the-life version.

We’ll walk through one realistic Azure scenario: a broken nightly pipeline, a silent schema change in a source system, and the conversation when the business wants answers. Along the way we’ll unpack what an Azure data engineer actually does to keep things running.


The Setup: A Very Normal Azure Data Platform

Our fictional environment is typical for a mid-sized analytics team:

  • Source systems
    • An operational SQL database for orders (Azure SQL Database)
    • A SaaS CRM exposed via REST APIs
  • Ingestion & orchestration
    • Azure Data Factory (ADF) with:
      • A nightly pipeline scheduled via a trigger
      • Copy activities pulling data from Azure SQL DB and REST APIs into Data Lake
      • A Data Flow transforming raw data to curated tables
  • Storage & modeling
    • Azure Data Lake Storage Gen2 (ADLS) as the main data lake
    • Lakehouse-style layout: /raw, /curated, /semantic
    • Azure Synapse Analytics serverless SQL for ad-hoc queries
    • Power BI datasets built on curated parquet files via dataflows

The key pipeline:

  1. Copy SalesOrders from Azure SQL DB to raw/sales_orders/YYYY/MM/DD/ (parquet)
  2. Copy CRM accounts from REST API to raw/accounts/
  3. Data Flow joins and cleans data into curated/sales/ partitioned by OrderDate
  4. Power BI dataflow refreshes from curated sales into a dataset

The business expectation: by 07:00, the “Daily Sales” dashboard shows complete data for the previous day.


07:00 – The Pipeline Failed Overnight

The first thing an Azure data engineer does most mornings is look for red.

Step 1: Check the ADF Monitor

You open the Monitor view in Azure Data Factory and see:

  • The nightly pipeline PL_Nightly_Sales failed at 02:13.
  • The failure is in the DF_Curated_Sales Data Flow activity.
  • Error message (typical example):

Column 'CustomerSegment' not found in input stream 'SalesOrders'.

This already hints at a schema change in the source.

Step 2: Confirm What Actually Ran

You drill into the pipeline run:

  • Copy from Azure SQL DB succeeded.
  • Copy from CRM API succeeded.
  • Data Flow failed.

So ingestion worked, transformation didn’t. That matters for how you respond to stakeholders: data is in the lake, but not yet in the curated layer the dashboard uses.

Step 3: Validate the Raw Data

You open the latest file in ADLS (e.g. via Azure Storage Explorer or Synapse serverless):

SELECT TOP 50 *
FROM OPENROWSET(
    BULK 'https://<storage-account>.dfs.core.windows.net/datalake/raw/sales_orders/2026/10/05/*.parquet',
    FORMAT = 'PARQUET'
) AS [r];

You confirm:

  • The data for yesterday is present.
  • The column CustomerSegment is indeed missing.

This is the first core task of an Azure data engineer: diagnose where the failure occurred and whether data is safely landed anywhere.


07:30 – The Hidden Schema Change

Nobody told you the source system changed. But the data flow was relying on a column that no longer exists.

Typical causes:

  • A developer removed or renamed a column in the Azure SQL DB.
  • A release script altered the table without updating downstream pipelines.

Step 4: Track the Schema Change at the Source

You query the source schema directly:

SELECT COLUMN_NAME, DATA_TYPE
FROM INFORMATION_SCHEMA.COLUMNS
WHERE TABLE_NAME = 'SalesOrders'
ORDER BY ORDINAL_POSITION;

You see CustomerSegment is gone, and a new column CustomerTier was added.

From an Azure platform perspective, nothing in ADF auto-adjusts for this. The Copy activity will happily copy whatever the source schema is; the Data Flow fails because its projection expects a column that no longer exists.

Step 5: Decide the Immediate Fix vs Proper Design

You have two competing pressures:

  • Get the dashboard green quickly.
  • Avoid hacking the pipeline into an inconsistent state.

You usually have three options:

  1. Hotfix the Data Flow to use the new column
    • Update the mapping to use CustomerTier instead of CustomerSegment.
    • Add logic to map CustomerTier into the semantic concept the report expects.
  2. Temporarily drop the column from the curated output
    • Remove CustomerSegment from the Data Flow output.
    • Adjust Power BI measures to not rely on it.
  3. Recreate the old column upstream (if feasible)
    • Add a computed column in the source (or in a staging view) that mimics CustomerSegment from CustomerTier.

What an Azure data engineer actually does all day is make this trade-off repeatedly: short-term continuity vs long-term model integrity.

In most real teams, the morning answer is: get the pipeline green with the least semantic damage, then schedule a proper model change.


08:00 – Making the Pipeline Green Again

Let’s say you choose option 1: adapt to CustomerTier.

Step 6: Update the Data Flow Projection

In the ADF Data Flow designer, you:

  • Open DF_Curated_Sales.
  • Refresh the source projection for SalesOrders so it picks up the new schema.
  • Replace references to CustomerSegment with CustomerTier.

You might add a derived column to keep the semantic name stable:

-- Expression in Data Flow derived column (using Data Flow expression language)

iif(CustomerTier == 'Gold', 'Premium',
    iif(CustomerTier == 'Silver', 'Standard', 'Other'))

That derived column is named CustomerSegment in the Data Flow output, so downstream tables and Power BI don’t break immediately.

Important detail: in Data Flows, projection refresh is manual. If you don’t refresh, the Data Flow still thinks CustomerSegment exists, and the run will keep failing.

Step 7: Re-run the Pipeline Safely

You trigger a manual rerun of the nightly pipeline for yesterday’s date:

  • Use a pipeline parameter pProcessingDate (string or date).
  • Manually set it to yesterday.

This is where design matters. A resilient Azure pipeline typically:

  • Drives file paths and filters from a parameter (e.g. 2026-10-05).
  • Is idempotent for that date: re-running doesn’t corrupt data.

You ensure the curated sink writes with overwrite semantics for that partition, or uses a delete-and-insert pattern.

After the rerun:

  • The pipeline succeeds.
  • Curated sales data for yesterday is present.

You now check the Power BI side.


08:30 – Power BI Is Still Late

Even with the pipeline fixed, the dashboard can still be wrong or stale. In this scenario, the Power BI dataset refresh depends on the curated layer.

Step 8: Check Power BI Refresh Status

In the Power BI Service:

  • Open the dataset linked to the Daily Sales report.
  • Check Refresh history.

You see:

  • A scheduled refresh at 06:30 failed.
  • Error message (example):

The key didn't match any rows in the table.

This often occurs when a schema change hits Power Query or the dataflow.

Step 9: Fix the Dataflow or Power Query

If you’re using Power BI dataflows on top of curated parquet:

  • Open the dataflow.
  • Check the Power Query steps.

Typical issues:

  • A Table.RenameColumns step referencing CustomerSegment throws an error.
  • A Table.SelectColumns step expects CustomerSegment but the column name changed.

You update the query to align with the new derived column in curated sales. For example:

let
    Source = AzureStorage.DataLake("https://<storage-account>.dfs.core.windows.net", [HierarchicalNavigation=true]),
    Curated = Source{[Name="datalake"]}[Data]{[Name="curated"]}[Data]{[Name="sales"]}[Data],
    Filtered = Table.SelectRows(Curated, each [OrderDate] >= Date.AddDays(Date.From(DateTime.LocalNow()), -1)),
    Selected = Table.SelectColumns(Filtered, {"OrderId", "OrderDate", "CustomerId", "CustomerSegment", "Revenue"})
in
    Selected

Here, CustomerSegment is still present because you kept the name stable in the Data Flow. If you had changed it to CustomerTier end-to-end, you’d adjust this step accordingly.

You then trigger a dataset refresh manually. Once it succeeds, the dashboard is live again.


09:00 – Explaining to Stakeholders Why the Dashboard Was Late

The third core part of the Azure data engineer role is translation: explaining technical causes and realistic guarantees to non-technical stakeholders.

You’ll typically cover three points:

  1. What broke
    • A schema change in the SalesOrders source removed a column the pipeline and report relied on.
    • The nightly transformation failed safely, stopping before writing partial or inconsistent data.
  2. What you’ve done
    • Adjusted the transformation to use the new source column and maintain the business concept.
    • Reran the pipeline for yesterday and fixed the Power BI refresh.
  3. What will change longer-term
    • Introduce schema-change detection and impact analysis.
    • Tighten the release process so source changes are communicated before deployment.

The details matter. For example, you can explain:

  • ADF doesn’t automatically adapt transformations to schema changes; that’s by design to avoid silent data drift.
  • The system is built to fail fast and visibly rather than silently produce wrong numbers.

This is often where the stakeholder finally understands what an Azure data engineer does beyond “building pipelines”.


10:00 – Hardening the Platform Against Next Time

Once the fire is out, the rest of the day is the real job: making sure this happens less often and is easier to fix.

1. Add Schema Drift Monitoring

You can implement checks such as:

  • A pre-transformation step that compares current source schema to a stored baseline.
  • A pipeline that queries INFORMATION_SCHEMA.COLUMNS or system views, writes schema snapshots into ADLS, and alerts on differences.

Example: a simple ADF pipeline that:

  1. Runs a Stored Procedure in Azure SQL DB that returns the current schema.
  2. Writes the result as a parquet file in metadata/schemas/salesorders/.
  3. A subsequent step compares the latest snapshot to yesterday’s and fails the nightly pipeline early if there’s an unexpected change.

This doesn’t prevent schema changes, but it moves the failure to a clear, early stage and makes impact analysis easier.

2. Make Pipelines Parameterized and Idempotent

To support reruns:

  • Use pipeline parameters for ProcessingDate.
  • Drive all file paths and filters from that parameter.
  • Configure sink writes so re-running for a given date either overwrites or safely merges.

A simple example: in a Copy activity to parquet, use a dynamic path:

@concat('raw/sales_orders/', formatDateTime(pProcessingDate, 'yyyy/MM/dd'), '/')

This makes correcting a failed run for a specific date straightforward.

3. Separate Raw, Curated, Semantic Layers Clearly

The scenario showed why layer separation matters:

  • Raw layer continued to ingest even when the transformation failed.
  • Curated layer failed, protecting semantic consumers from half-baked data.
  • Power BI sat on top of curated, not raw.

An Azure data engineer spends a lot of time enforcing this separation so that failures are localized and recoverable.

4. Improve Change Management with Source Teams

Some of the work is social, not technical:

  • Agree a change checklist with the team owning the Azure SQL DB.
  • Require a simple schema change announcement before deployment.
  • Maintain a shared data contract for key tables.

None of this is an Azure feature, but it’s central to the Azure data engineer role.


What Does an Azure Data Engineer Actually Do All Day?

If you strip away the tooling names, the core activities in this scenario are what define the Azure data engineer role:

  • Detect and diagnose failures in ADF pipelines, Data Flows, and Power BI refreshes.
  • Trace issues back to source changes, especially schema changes.
  • Design and maintain data models that survive operational changes.
  • Build and adjust pipelines that move data through raw, curated, and semantic layers.
  • Communicate impact and trade-offs to stakeholders who care about numbers, not columns.

The broken overnight pipeline, the schema change, and the late dashboard are not edge cases—they’re the normal rhythm of the job.


One Practical Takeaway

If you remember only one thing from this day-in-the-life: make yesterday’s data a parameter. When your Azure pipelines, storage paths, and Power BI queries all hinge on an explicit ProcessingDate, recovering from a broken run becomes a controlled, repeatable operation instead of a scramble through ad-hoc fixes.

Editor's Note

This article reflects the shift from ad-hoc pipeline building to disciplined, failure-aware data engineering in Azure environments where nightly loads and business-critical dashboards depend on fragile upstream schemas.

Professionals who want to apply these patterns to their own data can explore Excelgoodies' Data Engineering & BI Azure (On Cloud) programme - taught live by instructors, with certification awarded once a real project is running at work.

Insights compiled through ongoing industry research and discussions within the Excelgoodies Analytics Community.

Azure

New

Next Batches Now Live

Power BIPower BI
SQLSQL
Power AppsPower Apps
Power AutomatePower Automate
Microsoft FabricMicrosoft Fabrics
AzureAzure Data Engineering
Explore Dates & Reserve Your Spot → Reserve Your Spot →