Excelgoodies logo +31 97 010285556

LEARN THIS HANDS ON

Full Stack BI (On-Cloud)

. Live Online FILLING FAST
View all upcoming batches
Unity Catalog Explained: Why Every Databricks Job Ad Suddenly Mentions It

Unity Catalog Explained: Why Every Databricks Job Ad Suddenly Mentions It

You open a Dutch data engineering vacancy and there it is again: “3+ years of Unity Catalog experience” for a platform that only recently became the default. Meanwhile your Databricks workspace still runs on the old Hive metastore and a pile of manual ACLs. You’re wondering what exactly hiring managers expect when they say Unity Catalog Databricks and how it changes governance in practice.

This article walks through Unity Catalog from the viewpoint of a team migrating real workloads: how governance, permissions, lineage and the metastore fit together, what breaks, and what recruiters actually mean when they ask for Unity Catalog experience.

Scenario: The Lakehouse That Grew Too Fast

Picture a typical setup:

  • One Azure Databricks workspace shared by data engineers, analytics engineers and BI developers.
  • Data in Azure Data Lake Storage (ADLS) Gen2, mounted with dbutils.fs.mount.
  • Tables stored in the workspace’s Hive metastore, plus plenty of direct file access.
  • Power BI and other tools reading from SQL Warehouses and clusters.

Pain points:

  • No central view of who can access which datasets.
  • Multiple copies of the same data in different schemas.
  • Hard to answer basic questions: “Who used customer_data last week?” or “Which dashboard breaks if we change this schema?”

Management now wants:

  • Consistent governance across workspaces.
  • Fine‑grained permissions at table, column and row level.
  • Reliable lineage: which jobs feed which tables and reports.

That is exactly the problem space Unity Catalog was built for.

What Unity Catalog Actually Is (And Is Not)

Unity Catalog is Databricks’ unified governance layer for data and AI assets. At a high level it provides:

  • A central metastore shared across workspaces.
  • A three‑level namespace: catalog.schema.table.
  • Fine‑grained access control on catalogs, schemas, tables, views, functions, models and more.
  • Centralized connection to external storage (external locations).
  • Data lineage across notebooks, jobs, SQL queries, and downstream assets.
  • Row and column level security via dynamic views and column masking policies.

What it is not:

  • Not a separate compute engine; it governs what your clusters and SQL Warehouses can see.
  • Not a replacement for ADLS or other storage; it manages metadata and access to those locations.
  • Not optional long‑term: on current Databricks runtimes, Unity Catalog is the strategic path and many new features are UC‑only.

In our scenario, moving from the Hive metastore to Unity Catalog means:

  • Your old default.customer_data table becomes something like main.analytics.customer_data.
  • Permissions move from ad‑hoc cluster ACLs and storage mounts into Unity Catalog grants.
  • Jobs and SQL Warehouses use UC‑enabled compute so they can access the new namespace.

The Unity Catalog Metastore: The New Center of Gravity

From Hive Metastore to Unity Catalog Metastore

Previously, each Databricks workspace had its own Hive metastore. Tables were referenced as database.table, and sharing across workspaces meant copying data or using external tools.

Unity Catalog introduces:

  • Metastore: A governance boundary associated with a region and one or more workspaces.
  • Catalogs inside a metastore (e.g. main, finance, ml).
  • Schemas inside catalogs (e.g. main.analytics, finance.raw).

So the full name is:

SELECT *
FROM main.analytics.customer_data;

Key points:

  • A workspace is assigned to exactly one Unity Catalog metastore.
  • A metastore can be shared by multiple workspaces in the same region.
  • Tables can be managed (storage location controlled by Databricks) or external (data lives in your own storage locations).

In the scenario team, this means:

  • No more per‑workspace Hive metastores with diverging schemas.
  • One central metastore where all shared datasets live.
  • New workspaces (e.g. for dev, test) can attach to the same metastore, keeping governance consistent.

External Locations and Storage Credentials

Unity Catalog does not replace ADLS; it governs access to it.

Core concepts:

  • Storage credential: securely stored access to a storage account or container (e.g. via a managed identity or service principal).
  • External location: a named reference to a path in that storage (e.g. abfss://bronze@datalake.dfs.core.windows.net/).

You then create external tables pointing at data in those locations.

Example (SQL in Databricks):

CREATE EXTERNAL TABLE main.raw.customer_data
USING DELTA
LOCATION 'abfss://raw@datalake.dfs.core.windows.net/customer_data';

Once this is in Unity Catalog:

  • Access is controlled via GRANT statements on main.raw.customer_data and the external location.
  • You no longer need dbutils.fs.mount for shared datasets; UC handles the mapping.

Permissions: Why Governance Feels Different Under UC

Unity Catalog centralizes permissions at the object level instead of the workspace or cluster level.

Core Permission Model

You manage access with SQL GRANT/REVOKE on:

  • Catalogs and schemas: control who can create or browse objects.
  • Tables and views: control read/write access to data.
  • External locations: control who can create external tables or volumes.

Example: give the analytics team read‑only access to a schema:

GRANT USAGE ON CATALOG main TO `analytics_group`;
GRANT USAGE ON SCHEMA main.analytics TO `analytics_group`;
GRANT SELECT ON ALL TABLES IN SCHEMA main.analytics TO `analytics_group`;

You can also grant on future tables:

GRANT SELECT ON FUTURE TABLES IN SCHEMA main.analytics TO `analytics_group`;

That last line is one of the reasons hiring managers care about Unity Catalog experience: it changes how you design permission patterns.

Row‑Level and Column‑Level Security

Unity Catalog supports:

  • Row‑level security using views that filter based on user identity.
  • Column‑level security using column masking policies.

A simple row‑level security pattern for our scenario’s customer_data table:

CREATE OR REPLACE VIEW main.analytics.customer_data_rls AS
SELECT *
FROM main.analytics.customer_data
WHERE region IN (
  SELECT region
  FROM main.security.user_regions
  WHERE principal = current_user()
);

GRANT SELECT ON VIEW main.analytics.customer_data_rls TO `analytics_group`;

Now users see only rows for regions assigned to them.

Column masking uses policies defined in Unity Catalog and applied to columns; the logic is similar, but driven by policy objects rather than inline CASE expressions.

Identity and Groups

On Azure Databricks, Unity Catalog integrates with:

  • Workspace users and groups (often synchronized from Microsoft Entra ID / Azure AD).
  • Service principals for automation.

Best practice in the scenario team:

  • Map Entra ID / Azure AD groups (e.g. data-engineers, bi-developers) to Databricks groups.
  • Grant permissions to groups, not individuals.
  • Use service principals for jobs that need stable, non‑human access.

This is the governance shift recruiters are hinting at: you’re expected to design group‑based patterns, not ad‑hoc per‑user grants.

Lineage: Finally Seeing End‑to‑End Data Flows

Unity Catalog adds built‑in lineage for many operations executed on UC‑enabled compute:

  • Notebooks and jobs that read/write UC tables.
  • SQL queries and dashboards in Databricks SQL.
  • Delta Live Tables and other pipeline frameworks.

Lineage shows:

  • Which upstream tables feed a given table or view.
  • Which notebooks, jobs or dashboards depend on a table.
  • How data flows across catalogs and schemas.

In the scenario team, this answers questions like:

  • “If we drop column customer_segment, which jobs break?”
  • “Where does the data in main.analytics.sales_summary actually come from?”

Unity Catalog lineage is automatically captured when:

  • You use UC objects (catalog.schema.table) on UC‑enabled clusters/SQL Warehouses.
  • You use supported APIs and operations (e.g. standard Spark SQL, Delta operations).

It does not retroactively cover non‑UC tables or direct file access. That’s why part of the migration effort is moving jobs from file paths to UC tables.

Migrating the Scenario Team: What Actually Changes

1. Enable Unity Catalog and Attach Workspaces

On Azure Databricks, Unity Catalog is configured at the account level:

  • Create a metastore in the region where your workspaces run.
  • Configure a default storage location for managed tables.
  • Assign your workspaces to that metastore.

Once a workspace is attached, UC‑enabled compute can access catalogs and schemas in that metastore.

2. Move from Mounts to External Locations

In our scenario, the team has many mounts like:

# Existing pattern (Hive + mounts)
dbutils.fs.mount(
    source="abfss://raw@datalake.dfs.core.windows.net/",
    mount_point="/mnt/raw",
    extra_configs={"fs.azure.account.auth.type": "OAuth", ...}
)

Under Unity Catalog, the recommended pattern is:

  • Define storage credentials and external locations in UC.
  • Create external tables pointing to those locations.
  • Stop using dbutils.fs.mount for shared datasets.

This centralizes access control and makes lineage and permissions consistent.

3. Create Catalogs and Schemas That Reflect Domains

Instead of one giant default database, you can now structure data by domain:

  • main.raw for raw ingested data.
  • main.curated for cleaned, modeled tables.
  • finance.reporting, marketing.analytics, etc.

Example:

CREATE CATALOG main;
CREATE SCHEMA main.raw;
CREATE SCHEMA main.curated;

Then migrate tables:

CREATE TABLE main.curated.customer_data
AS SELECT * FROM default.customer_data;

(Actual migration may use CREATE TABLE ... LOCATION or ALTER TABLE SET LOCATION to avoid copying data; choose based on your current storage layout.)

4. Switch Jobs to UC‑Enabled Compute and Namespaces

Jobs and clusters need to run in Unity Catalog‑enabled mode to access UC objects. On Azure Databricks this means:

  • Use clusters or SQL Warehouses that have UC turned on.
  • Reference tables with full catalog.schema.table names.

Example job SQL before:

INSERT INTO analytics.customer_data_clean
SELECT * FROM customer_data_raw;

After UC migration:

INSERT INTO main.curated.customer_data_clean
SELECT *
FROM main.raw.customer_data_raw;

Once jobs run on UC‑enabled compute and use UC tables, lineage and governance apply automatically.

5. Apply Governance Patterns Instead of One‑Off Fixes

With UC in place, the team can implement:

  • Domain‑based access: BI developers get SELECT on curated schemas, data engineers get ALL PRIVILEGES on raw and curated.
  • Future grants: new tables inherit the right permissions without manual GRANT.
  • RLS and masking: sensitive tables exposed via secured views.

This is the kind of experience job ads are hinting at: not just “clicked around in Unity Catalog”, but “designed and operated governance patterns on UC”.

Why Job Ads Ask for “Years of Unity Catalog Experience”

Even though Unity Catalog is relatively new, recruiters and hiring managers use “Unity Catalog experience” as shorthand for several skills:

  • Designing catalog/schema structures that match business domains and lifecycle (raw/curated/marts).
  • Migrating from Hive metastore and mounts to UC tables and external locations without breaking workloads.
  • Implementing permission models with groups, future grants, RLS and column masking.
  • Operating multi‑workspace environments on a shared metastore.
  • Troubleshooting UC‑specific issues: missing privileges, non‑UC clusters, lineage gaps, external location misconfigurations.

If you can speak concretely about these areas, you effectively have the “Unity Catalog Databricks governance” experience those ads are asking for, regardless of how many calendar years UC has existed.

One Practical Takeaway for Your Next Week at Work

If your team is still on the Hive metastore, pick one non‑critical dataset and:

  1. Create a catalog and schema in Unity Catalog for it.
  2. Define an external location pointing at its ADLS path.
  3. Create an external table in UC.
  4. Grant access to a single group.
  5. Switch one job or notebook to read/write that UC table on UC‑enabled compute.

You’ll immediately see how governance, permissions and lineage behave under Unity Catalog, and you’ll have a concrete story to tell the next time a job ad asks about Databricks governance.

Editor's Note

This article reflects the shift from workspace-level Hive metastore setups to shared Unity Catalog governance in Databricks, especially for teams consolidating Azure lakehouse environments and tightening access control across domains.

Professionals who want to apply these patterns to their own data can explore Excelgoodies' Data Engineering & BI Azure (On Cloud) programme - taught live by instructors, with certification awarded once a real project is running at work.

Insights compiled through ongoing industry research and discussions within the Excelgoodies Analytics Community.

Azure

New

Next Batches Now Live

Power BIPower BI
SQLSQL
Power AppsPower Apps
Power AutomatePower Automate
Microsoft FabricMicrosoft Fabrics
AzureAzure Data Engineering
Explore Dates & Reserve Your Spot → Reserve Your Spot →