Lakehouse Health Intelligence Accelerator scans storage, compute, jobs and security, scores their health, and hands you a prioritized, costed fix list. It runs inside your own workspace and reads only metadata, never your data.
Lakehouse Health Intelligence Accelerator
See Your Entire Databricks Platform’s Health in One Score.
Automated Checks in Every Scan
Assessment Engines: Storage, Compute, Jobs, Security
Weighted Health Score by Pillar, Engine and Overall
Rows of Your Business Data Read or Moved
Operational Debt Builds Up Where Nobody Is Looking.
Every Databricks platform accumulates invisible debt as it grows. It's scattered across four domains, hard to see, and there is no single place to measure it.
Each issue is small. Together they burn cost, slow teams down, and create real risk.
-
Storageclutter
Tables with no owner or description, unoptimized and duplicated tables, small-file problems, and stale data nobody has touched in a year.
-
Computewaste
Clusters that never auto-terminate, idle SQL warehouses, missing cluster policies, and interactive compute running scheduled jobs.
-
Jobsfragility
Jobs with no alerting, jobs pinned to personal notebooks, missing tags, and hundreds of stale notebooks.
-
Securityexposure
Over-broad grants, public access, blanket schema-level permissions, and unrestricted external sharing.
The result is higher cost, slower teams and compliance risk, with no easy way to measure or improve it. Lakehouse Health Intelligence makes it visible, measurable and actionable.
The Operational Health and Optimization Layer for Databricks.
A configuration-driven assessment built from native Databricks notebooks. No external service and nothing to send data to. Everything it produces lives in a Unity Catalog in your workspace, owned and controlled by you.
health_intelligenceconfiginventoryhealthlogs
Four Engines. 115 Checks One View of the Whole Platform.
Each engine reads a different slice of platform metadata and groups its checks into pillars, so findings roll up into capability areas rather than a flat list.
24
Storage Checks
Reads table and column metadata from the information schema, plus DESCRIBE DETAIL and DESCRIBE HISTORY. It sees table sizes, file counts, formats, owners, descriptions, partitioning, and OPTIMIZE and VACUUM history.
Pillars
- Governance
- Performance
- Optimization
- Cost
- Lifecycle
- Reliability
- Duplication
Example Checks
- Missing table owner
- Missing table description
- Missing column descriptions
- Small file problem
- Large table missing OPTIMIZE
- Stale tables
- Tiny tables
- Clone and backup detection
28
Compute Checks
Reads every cluster and SQL warehouse configuration plus query history: auto-termination, autoscaling, Photon, spot usage and idle time.
Focus Areas
- Configuration
- Cost
- Reliability
Example Checks
- Missing cluster policy
- No auto-termination
- Autoscaling disabled
- Photon disabled
- Idle SQL warehouse
- Interactive compute used for scheduled jobs
- Unapproved runtime version
35
Jobs and Orchestration Checks
Reads all jobs, pipelines and notebooks: schedules, alerting, tags, notebook locations, run history and staleness.
Focus Areas
- Governance
- Reliability
- Cost
- Operations
Example Checks
- Jobs without alerting
- Jobs using personal workspace notebooks
- Notebook without Git integration
- Development job on schedule
- Missing required tags
- Stale notebooks
- Long-running jobs
- Oversized job cluster
28
Security and Governance Checks
Reads grants, users, service principals, secret scopes, external locations, storage credentials, Delta shares and audit events.
Pillars
- Access control
- Authentication
- Classification
- External sharing
- Lineage
- Secrets
- Audit
Example Checks
- Tables with public access
- Grants to ACCOUNT USERS
- Schema-level blanket grants
- External locations without IP restrictions
- Storage credentials with broad access
- Unused external locations
- Sensitive tables shared externally
Every source is metadata: sizes, configs, permissions and history. It never runs a SELECT on your business tables, and your data never moves.
One Run, Six Stages, from Inventory to Action Plan.
A single orchestrator notebook drives the whole pipeline. Run it on demand or put it on a schedule.
Setup
On first run it creates its own Unity Catalog,
health_intelligence, with four schemas: config, inventory, health and logs.Discovery
Four collectors run in parallel and write platform metadata into about 20 inventory tables, from table and cluster inventory to security grants and Delta shares.
Assessment
The four engines evaluate all 115 checks and record every finding with its severity, the exact asset, the issue and a recommendation.
Scoring
A weighted 0–100 health score per pillar, engine and overall, weighted by severity and by how many objects each check affects.
Recommendations and Cost
Findings become per-asset fixes priced at Databricks list rates for storage reclaim and DBU savings, with ready-to-run remediation SQL where the fix is deterministic.
Change Tracking and Dashboard
Each scan is compared with the last one as new, resolved or still open, and the five-tab dashboard is published.
Weighted for real blast radius. A failing check that touches 50 tables counts for more than one that touches a single table, so the score is defensible rather than a vanity number.
Watch the Walkthrough
See how one scan turns 115 automated checks into a health score, a prioritized fix list and a costed plan for your Databricks platform.
Your Definition of Healthy, Not Ours.
Different businesses hold different standards. Every threshold, every check's on/off switch and the scan scope live in a simple CSV and JSON config you can edit without touching code.
- Retune a threshold, for example stale tables at 90 days instead of 365
- Turn individual checks on or off
- The file is validated and applied on every run
threshold_config.csv
| Check | Threshold | Enabled |
|---|---|---|
| Stale table | 365 days | true |
| Small file problem | 128 files | true |
| Tiny tables | 100 MB | true |
| Idle SQL warehouse | 7 days | true |
| Large table missing OPTIMIZE | 30 days | true |
Default thresholds shown. Edit the threshold and enabled columns to match your standards.
A View for Leadership, and a Deep Dive for Every Team.
The five-tab dashboard opens on an executive overview, then gives Storage, Compute, Jobs and Security their own tab with the same layout, all from one scan.
Assets evaluated, findings by severity, cost, storage and DBU opportunity, and the overall health score with its trend across scans.
Where risk concentrates by engine and severity, with an engine summary of checks passed, failed and status.
The exact asset, the failing check and what's wrong. A ready-made to-do list: start at the top.
Each engine tab scopes the KPIs, severity counts, cost opportunity and health gauge to that engine alone.
Pinpoints the weak area. Here, Storage governance scores 41.6 while every other pillar sits between 84 and 100.
Built to Pass Security Review and Keep Paying Off
Metadata Only
Reads platform metadata, never table rows. Nothing leaves your tenant, which means easier security approval and no PII exposure.
In-Tenant and Yours
Installs natively. All results live in your own
health_intelligencecatalog, queryable and extensible.Fully Configurable
Thresholds, check switches and scan scope are yours to edit, so the assessment adapts to your organization.
Actionable, Not Just Diagnostic
Prioritized fixes with dollar impact and ready-to-run remediation SQL where the fix is deterministic.
Repeatable and Measurable
Designed to run on a schedule and prove health is improving, scan over scan.
Whole-Platform Breadth
Storage, compute, jobs and security in one score, in one place.
What It Changes for Each Team.
One scan produces a view for the operators who fix things, the teams who watch cost and risk, and the leaders who track progress.
Platform Admins and Data Engineering Leads
Find unowned and undocumented assets, missing policies and tags, fragile jobs, missing alerting and stale assets, with a prioritized list to work through.
Stronger Governance and ReliabilityFinOps
Surface idle warehouses, oversized clusters, reclaimable storage and wasted DBUs, each with a quantified savings opportunity.
Lower CostSecurity and Governance
Expose over-broad access, public grants and unrestricted external sharing before they become incidents, alongside your audit posture.
Reduced RiskEngineering Managers and CTOs
One measurable health score leadership can track, and a trend that shows whether the platform is getting better.
Measurable Maturity
At a Glance
Engines
4: Storage, Compute, Jobs, Security
Data accessed
Metadata only, never business table contents
Output
health_intelligence catalog with 30+ tables, and a five-tab dashboard
Score
Weighted 0–100 per pillar, engine and overall
Cadence
On demand or scheduled
Checks
115 (Storage 24, Compute 28, Jobs 35, Security 28)
Runs in
Your own Databricks workspace
Deliverables
Health score, prioritized findings, cost opportunity, remediation SQL, trend
Configurable
Thresholds, enable or disable, scan scope
Frequently Asked Questions
Lakehouse Health Intelligence Accelerator is a Databricks health check and optimization accelerator. It runs 115 automated checks across storage, compute, jobs and security, produces a weighted 0–100 health score, and delivers a prioritized fix list with estimated cost savings. It runs inside your own Databricks workspace and reads only metadata.
Lakehouse Health Intelligence Accelerator installs as native Databricks notebooks. You run a single orchestrator notebook, which creates a health_intelligence catalog in Unity Catalog, collects platform metadata, evaluates all 115 checks, scores the results and publishes a five-tab dashboard. Run it on demand or on a schedule to track improvement scan over scan.
Start by finding compute and storage waste. Lakehouse Health Intelligence Accelerator flags idle SQL warehouses, clusters without auto-termination, disabled autoscaling, oversized job clusters, interactive compute used for scheduled jobs, and reclaimable storage. Each finding is priced using Databricks list rates for DBU and storage savings, so you can fix the most valuable issues first.
Not with Lakehouse Health Intelligence Accelerator. It reads only platform metadata such as table sizes, cluster configurations, permissions and run history. It never runs a SELECT on your business tables, and no data leaves your tenant, which simplifies security review and avoids PII exposure.
Start with visibility into who can access what. The Security and Governance engine checks access control, authentication, external sharing, secrets and audit, and flags issues like tables with public access, grants to ACCOUNT USERS, schema-level blanket grants and sensitive tables shared externally. The Storage engine adds checks for tables missing owners and descriptions.
The health score is a weighted 0–100 value, weighted by each finding's severity and by how many objects each check affects, so it reflects real impact rather than a simple pass rate. You get a score per pillar, per engine and for the whole platform, based on thresholds you can configure to match your own standards.
Turn Invisible Operational Debt into a Costed Plan.
See Lakehouse Health Intelligence Accelerator run end to end, and find out what a scan of your own platform would surface.