Lakehouse Health Intelligence Accelerator

See Your Entire Databricks Platform’s Health in One Score.

Automated Checks in Every Scan
Assessment Engines: Storage, Compute, Jobs, Security
Weighted Health Score by Pillar, Engine and Overall
Rows of Your Business Data Read or Moved

Lakehouse Health Intelligence Accelerator scans storage, compute, jobs and security, scores their health, and hands you a prioritized, costed fix list. It runs inside your own workspace and reads only metadata, never your data.

Operational Debt Builds Up Where Nobody Is Looking.

Every Databricks platform accumulates invisible debt as it grows. It's scattered across four domains, hard to see, and there is no single place to measure it.

Each issue is small. Together they burn cost, slow teams down, and create real risk.

  • Storageclutter

    Tables with no owner or description, unoptimized and duplicated tables, small-file problems, and stale data nobody has touched in a year.

  • Computewaste

    Clusters that never auto-terminate, idle SQL warehouses, missing cluster policies, and interactive compute running scheduled jobs.

  • Jobsfragility

    Jobs with no alerting, jobs pinned to personal notebooks, missing tags, and hundreds of stale notebooks.

  • Securityexposure

    Over-broad grants, public access, blanket schema-level permissions, and unrestricted external sharing.

The result is higher cost, slower teams and compliance risk, with no easy way to measure or improve it. Lakehouse Health Intelligence makes it visible, measurable and actionable.

The Operational Health and Optimization Layer for Databricks.

A configuration-driven assessment built from native Databricks notebooks. No external service and nothing to send data to. Everything it produces lives in a Unity Catalog in your workspace, owned and controlled by you.

health_intelligenceconfiginventoryhealthlogs

Four Engines. 115 Checks One View of the Whole Platform.

Each engine reads a different slice of platform metadata and groups its checks into pillars, so findings roll up into capability areas rather than a flat list.

24

Storage Checks

Reads table and column metadata from the information schema, plus DESCRIBE DETAIL and DESCRIBE HISTORY. It sees table sizes, file counts, formats, owners, descriptions, partitioning, and OPTIMIZE and VACUUM history.

Pillars

  • Governance
  • Performance
  • Optimization
  • Cost
  • Lifecycle
  • Reliability
  • Duplication

Example Checks

  • Missing table owner
  • Missing table description
  • Missing column descriptions
  • Small file problem
  • Large table missing OPTIMIZE
  • Stale tables
  • Tiny tables
  • Clone and backup detection

Every source is metadata: sizes, configs, permissions and history. It never runs a SELECT on your business tables, and your data never moves.

One Run, Six Stages, from Inventory to Action Plan.

A single orchestrator notebook drives the whole pipeline. Run it on demand or put it on a schedule.

  1. Setup

    On first run it creates its own Unity Catalog, health_intelligence, with four schemas: config, inventory, health and logs.

  2. Discovery

    Four collectors run in parallel and write platform metadata into about 20 inventory tables, from table and cluster inventory to security grants and Delta shares.

  3. Assessment

    The four engines evaluate all 115 checks and record every finding with its severity, the exact asset, the issue and a recommendation.

  4. Scoring

    A weighted 0–100 health score per pillar, engine and overall, weighted by severity and by how many objects each check affects.

  5. Recommendations and Cost

    Findings become per-asset fixes priced at Databricks list rates for storage reclaim and DBU savings, with ready-to-run remediation SQL where the fix is deterministic.

  6. Change Tracking and Dashboard

    Each scan is compared with the last one as new, resolved or still open, and the five-tab dashboard is published.

Weighted for real blast radius. A failing check that touches 50 tables counts for more than one that touches a single table, so the score is defensible rather than a vanity number.

Watch the Walkthrough

See how one scan turns 115 automated checks into a health score, a prioritized fix list and a costed plan for your Databricks platform.

Your Definition of Healthy, Not Ours.

Different businesses hold different standards. Every threshold, every check's on/off switch and the scan scope live in a simple CSV and JSON config you can edit without touching code.

  • Retune a threshold, for example stale tables at 90 days instead of 365
  • Turn individual checks on or off
  • The file is validated and applied on every run
threshold_config.csv
CheckThresholdEnabled
Stale table365 daystrue
Small file problem128 filestrue
Tiny tables100 MBtrue
Idle SQL warehouse7 daystrue
Large table missing OPTIMIZE30 daystrue

Default thresholds shown. Edit the threshold and enabled columns to match your standards.

A View for Leadership, and a Deep Dive for Every Team.

The five-tab dashboard opens on an executive overview, then gives Storage, Compute, Jobs and Security their own tab with the same layout, all from one scan.

Assets evaluated, findings by severity, cost, storage and DBU opportunity, and the overall health score with its trend across scans.

Executive overview tab with KPI tiles for assets evaluated and findings by severity, and an overall health score gauge at 89.3 percent.Overall health 89.3%, rated Excellent, across 573 assets.

Where risk concentrates by engine and severity, with an engine summary of checks passed, failed and status.

Stacked bar chart of findings by engine and severity next to a bar chart of health score by engine: Compute 92.5, Jobs 85, Security 94.2, Storage 85.7, with an engine summary table below.Jobs and Storage stand out as the areas that need attention first.

The exact asset, the failing check and what's wrong. A ready-made to-do list: start at the top.

Table of top assets requiring attention listing notebooks flagged by the Stale Notebook check, with days since last modified.Stale notebooks flagged with the number of days since they were last modified.

Each engine tab scopes the KPIs, severity counts, cost opportunity and health gauge to that engine alone.

Storage tab showing 85 assets evaluated, findings by severity, storage reclaim, and a Storage health score of 85.7 percent.The Storage tab: 85 assets evaluated and a health score of 85.7%.

Pinpoints the weak area. Here, Storage governance scores 41.6 while every other pillar sits between 84 and 100.

Storage count by severity donut chart and a health score by pillar bar chart where Governance scores 41.6, with an object evaluation table below.The problem is governance, meaning missing owners and descriptions, not performance or cost.

Every check by name and ID with objects evaluated, failed and passed, plus the raw findings for each asset.

Object evaluation table of storage checks with evaluated, failed and passed counts, above a findings table listing severity, check ID, asset and issue.Concrete and measurable: for example, 53 of 85 tables are missing a description.

Each finding carries its severity, the asset, the issue and a recommendation your team can act on.

Findings detail table listing jobs with no description or missing required tags, each with a recommendation.Jobs with no description or missing tags, each paired with a recommendation.

Built to Pass Security Review and Keep Paying Off

  • Metadata Only

    Reads platform metadata, never table rows. Nothing leaves your tenant, which means easier security approval and no PII exposure.

  • In-Tenant and Yours

    Installs natively. All results live in your own health_intelligence catalog, queryable and extensible.

  • Fully Configurable

    Thresholds, check switches and scan scope are yours to edit, so the assessment adapts to your organization.

  • Actionable, Not Just Diagnostic

    Prioritized fixes with dollar impact and ready-to-run remediation SQL where the fix is deterministic.

  • Repeatable and Measurable

    Designed to run on a schedule and prove health is improving, scan over scan.

  • Whole-Platform Breadth

    Storage, compute, jobs and security in one score, in one place.

What It Changes for Each Team.

One scan produces a view for the operators who fix things, the teams who watch cost and risk, and the leaders who track progress.

  • Platform Admins and Data Engineering Leads

    Find unowned and undocumented assets, missing policies and tags, fragile jobs, missing alerting and stale assets, with a prioritized list to work through.

    Stronger Governance and Reliability
  • FinOps

    Surface idle warehouses, oversized clusters, reclaimable storage and wasted DBUs, each with a quantified savings opportunity.

    Lower Cost
  • Security and Governance

    Expose over-broad access, public grants and unrestricted external sharing before they become incidents, alongside your audit posture.

    Reduced Risk
  • Engineering Managers and CTOs

    One measurable health score leadership can track, and a trend that shows whether the platform is getting better.

    Measurable Maturity

At a Glance

Engines

4: Storage, Compute, Jobs, Security

Data accessed

Metadata only, never business table contents

Output

health_intelligence catalog with 30+ tables, and a five-tab dashboard

Score

Weighted 0–100 per pillar, engine and overall

Cadence

On demand or scheduled

Checks

115 (Storage 24, Compute 28, Jobs 35, Security 28)

Runs in

Your own Databricks workspace

Deliverables

Health score, prioritized findings, cost opportunity, remediation SQL, trend

Configurable

Thresholds, enable or disable, scan scope

Frequently Asked Questions

Lakehouse Health Intelligence Accelerator is a Databricks health check and optimization accelerator. It runs 115 automated checks across storage, compute, jobs and security, produces a weighted 0–100 health score, and delivers a prioritized fix list with estimated cost savings. It runs inside your own Databricks workspace and reads only metadata.

Lakehouse Health Intelligence Accelerator installs as native Databricks notebooks. You run a single orchestrator notebook, which creates a health_intelligence catalog in Unity Catalog, collects platform metadata, evaluates all 115 checks, scores the results and publishes a five-tab dashboard. Run it on demand or on a schedule to track improvement scan over scan.

Start by finding compute and storage waste. Lakehouse Health Intelligence Accelerator flags idle SQL warehouses, clusters without auto-termination, disabled autoscaling, oversized job clusters, interactive compute used for scheduled jobs, and reclaimable storage. Each finding is priced using Databricks list rates for DBU and storage savings, so you can fix the most valuable issues first.

Not with Lakehouse Health Intelligence Accelerator. It reads only platform metadata such as table sizes, cluster configurations, permissions and run history. It never runs a SELECT on your business tables, and no data leaves your tenant, which simplifies security review and avoids PII exposure.

Start with visibility into who can access what. The Security and Governance engine checks access control, authentication, external sharing, secrets and audit, and flags issues like tables with public access, grants to ACCOUNT USERS, schema-level blanket grants and sensitive tables shared externally. The Storage engine adds checks for tables missing owners and descriptions.

The health score is a weighted 0–100 value, weighted by each finding's severity and by how many objects each check affects, so it reflects real impact rather than a simple pass rate. You get a score per pillar, per engine and for the whole platform, based on thresholds you can configure to match your own standards.

Turn Invisible Operational Debt into a Costed Plan.

See Lakehouse Health Intelligence Accelerator run end to end, and find out what a scan of your own platform would surface.