Business Professionals
Power BI | Power Pivot | Power Query | DAX
Cloud Flows | RPA | AI Builder | Copilot
60+ Formulas | Data Stories | Advanced Reporting & Modeling
VB Programming | Report Automation |
MS-Office Automation
Techno-Business Professionals
Power BI | Power Query | Advanced DAX | SQL - Query &
Programming
Microsoft Fabric | Power BI | Power Query | Advanced DAX |
SQL - Query & Programming
Power BI | Power Apps | Power Automate | Copilot Studio | Power Pages | Dataverse
Microsoft Power Apps | Microsoft Power Automate
Power BI | Adv. DAX | SQL (Query & Programming) |
VBA | Python | Web Scrapping | API Integration
Power BI | Power Apps | Power Automate |
SQL (Query & Programming)
Power BI | Adv. DAX | Power Apps | Power Automate |
SQL (Query & Programming) | VBA | Python | Web Scrapping | API Integration
Power Apps | Power Automate | SQL | VBA | Python |
Web Scraping | RPA | API Integration
Technology Professionals
Power BI | DAX | SQL | ETL with SSIS | SSAS | VBA | Python
Power BI | SQL | Azure Data Lake | Synapse Analytics |
Data Factory | Databricks | Power Apps | Power Automate |
Azure Analysis Services
Microsoft Fabric | Power BI | SQL | Lakehouse |
Data Factory (Pipelines) | Dataflows Gen2 | KQL | Delta Tables | Power Apps | Power Automate
Power BI | Power Apps | Power Automate | SQL | VBA | Python | API Integration
Power BI | Advanced DAX | Databricks | SQL | Lakehouse Architecture
Business Professionals
Power BI | Power Pivot | Power Query | DAX
Cloud Flows | RPA | AI Builder | Copilot
60+ Formulas | Data Stories | Advanced Reporting & Modeling
VB Programming | Report Automation |
MS-Office Automation
Techno-Business Professionals
Power BI | Power Query | Advanced DAX | SQL - Query &
Programming
Microsoft Fabric | Power BI | Power Query | Advanced DAX |
SQL - Query & Programming
Power BI | Power Apps | Power Automate | Copilot Studio | Power Pages | Dataverse
Microsoft Power Apps | Microsoft Power Automate
Power BI | Adv. DAX | SQL (Query & Programming) |
VBA | Web Scrapping | API Integration
Power BI | Power Apps | Power Automate |
SQL (Query & Programming)
Power BI | Adv. DAX | Power Apps | Power Automate |
SQL (Query & Programming) | VBA | Web Scrapping | API Integration
Power Apps | Power Automate | SQL | VBA |
Web Scraping | RPA | API Integration
Technology Professionals
Power BI | DAX | SQL | ETL with SSIS | SSAS | VBA
Power BI | SQL | Azure Data Lake | Synapse Analytics |
Data Factory | Azure Analysis Services
Microsoft Fabric | Power BI | SQL | Lakehouse |
Data Factory (Pipelines) | Dataflows Gen2 | KQL | Delta Tables
Power BI | Power Apps | Power Automate | SQL | VBA | API Integration
Power BI | Advanced DAX | Databricks | SQL | Lakehouse Architecture
A lot of BI developers in Utrecht and elsewhere hit the same wall: they can bend DAX to their will, shape data with Power Query, and ship polished reports—but freeze when someone mentions "Delta Lake" or "PySpark pipelines". If you’re wondering how to turn your Power BI developer experience into a credible Azure data engineer career path without starting from zero, this article gives you a concrete map.
We’ll follow one realistic scenario: a BI analyst who has built a successful semantic model on top of a messy source system, and whose team now wants a proper lakehouse and reusable data pipelines. We’ll walk the path from SQL and Power Query to Azure Data Factory, Spark, and lakehouses, showing what actually changes and what carries over.
You’re a BI developer in a mid-size organisation:
The pain points:
Your manager’s ask: “Can you help us design the lakehouse and pipelines? You already understand the data best.”
This is the BI developer to data engineer moment.
Before diving into tools, it helps to map your current skills to data engineering responsibilities.
From Power BI / BI development:
These are directly valuable in data engineering design decisions.
For an Azure data engineer role, the big gaps usually are:
You don’t need to become a full-time software engineer, but you do need enough depth to design and maintain production-grade data flows.
Start from something you already know: your existing Power BI dataset.
Take your main dataset and classify tables:
In Power BI, you probably do Bronze→Gold transformations in Power Query and sometimes in DAX. In a lakehouse, the goal is:
Pick one existing Power BI model and:
This gives you a first-pass design of your lakehouse layers without touching any new tools.
Power Query refresh is essentially a pipeline hidden inside Power BI. As volumes grow, you need explicit orchestration.
On Azure today, two common options are:
Power BI refresh:
ADF / Fabric Data Factory:
Take a typical Power Query flow:
In ADF or Fabric Data Factory, the analogous pipeline might:
You’re not losing Power Query skills here—you’re just moving the logic into a more scalable, observable environment.
The most intimidating piece for many BI developers is Spark. The good news: your SQL and Power Query experience translates better than you expect.
On Azure, you’ll typically encounter Spark in:
Spark:
DataFrames) across a cluster.Power Query:
The core operations—select, filter, join, group—are the same.
Suppose you have a CSV of sales in ADLS and a dimension of products in a SQL database. You want a clean Gold fact table in a lakehouse.
A basic PySpark notebook in Fabric or Databricks might look like this:
from pyspark.sql import SparkSession
from pyspark.sql.functions import col, to_date
spark = SparkSession.builder.getOrCreate()
# Bronze: read raw sales from ADLS CSV
sales_bronze = spark.read.option("header", "true").csv("abfss://data@storageaccount.dfs.core.windows.net/bronze/sales/")
# Silver: clean types and remove duplicates
sales_silver = (
sales_bronze
.withColumn("SaleDate", to_date(col("SaleDate"), "yyyy-MM-dd"))
.withColumn("Quantity", col("Quantity").cast("int"))
.withColumn("Amount", col("Amount").cast("decimal(18,2)"))
.dropDuplicates(["SaleId"])
)
# Read product dimension from a SQL database via JDBC
products_dim = spark.read.format("jdbc").options(
url="jdbc:sqlserver://myserver.database.windows.net:1433;database=mydb",
dbtable="dbo.Products",
user="sql_user",
password="sql_password",
driver="com.microsoft.sqlserver.jdbc.SQLServerDriver"
).load()
# Gold: join sales to products
sales_gold = (
sales_silver.alias("s")
.join(products_dim.alias("p"), col("s.ProductId") == col("p.ProductId"), "left")
)
# Write out as Delta Lake for lakehouse consumption
sales_gold.write.format("delta").mode("overwrite").save("abfss://data@storageaccount.dfs.core.windows.net/gold/sales/")
This does exactly what the surrounding text describes:
If you can read Power Query’s step list, you can reason about this PySpark pipeline.
“Lakehouse” is a broad term, but there are a few concepts that matter immediately when you move from Power BI to data engineering.
For production lakehouses on Azure, you’ll usually prefer:
As a BI developer, the key implications:
In Power BI, schema drift often shows up as a refresh error. In a lakehouse:
For example, in Delta, an unexpected new column can either be rejected or handled depending on options you set when writing.
In a modern Azure/Fabric stack you might have:
The data engineer career path is about owning the lakehouse and warehouse layers, not just the dataset.
As soon as you leave “Power BI Desktop + scheduled refresh”, you hit operational concerns.
Data engineering work typically lives in:
Practical steps for a BI developer:
Instead of “refresh failed, check gateway”, you’ll deal with:
Core habits:
You already manage row-level security in Power BI. Data engineering adds:
You don’t need to be the security architect, but you must design pipelines that respect least-privilege and avoid embedding credentials in code.
To make this actionable, here’s a practical sequence that fits into normal project work.
Document your existing model as Bronze/Silver/Gold.
Prototype a lakehouse on a non-critical subject area.
Rewrite a Power Query transformation in Spark.
Expose Gold tables to Power BI from the lakehouse.
Introduce basic DevOps practices.
Each step builds data engineering skills on top of your existing BI understanding rather than replacing it.
The fastest way to move from Power BI developer to data engineer is not to start with generic Spark tutorials—it’s to re-implement one of your existing, successful BI models as a lakehouse with pipelines. If you can turn a single dataset’s Power Query and DAX into Bronze/Silver/Gold tables, orchestrated with ADF or Fabric and processed with PySpark, you’ve crossed the real skills gap that nobody talks about—and you’ve done it in a way your team can put into production.
This article reflects the shift from report-centric BI development to lakehouse-oriented data engineering, especially in teams where Power BI practitioners are being asked to own Azure-based pipelines and shared data platforms.
Professionals who want to apply these patterns to their own data can explore Excelgoodies' Data Engineering & BI Azure (On Cloud) programme - taught live by instructors, with certification awarded once a real project is running at work.
Insights compiled through ongoing industry research and discussions within the Excelgoodies Analytics Community.
Azure
New
Next Batches Now Live
Power BI
SQL
Power Apps
Power Automate
Microsoft Fabrics
Azure Data Engineering