Data Quality Framework Accelerator
Catch Bad Data Before It Reaches the Business.
| Emp_ID | Emp_Name | DOB | Experience | Status | Validation remark | |
|---|---|---|---|---|---|---|
| E1001 | Priya Nair | priya.nair@corp.com | 1990-04-12 | 8 | ·P | Passed all rules |
| E1002 | Arjun Mehta | arjun.mehta@corp | 1987-11-03 | 12 | ·R | Invalid email |
| E1003 | null | l.fernandes@corp.com | 1995-02-27 | 4 | ·R | Emp_Name cannot be null or empty |
| E1004 | Sara Khan | sara.khan@corp.com | 1994-13-40 | 6 | ·R | DOB must be a valid date |
| E1005 | Rohit Das | rohit.das@corp.com | 1992-07-19 | five | ·R | Experience must be integer |
| E1001 | Priya Nair | priya.nair@corp.com | 1990-04-12 | 8 | ·D | Duplicate record |
Every Dataset Breaks in Its Own Way
By the time a dashboard looks wrong, the bad record has already travelled through reports and business processes. These are the issues that cause it.
- Required fields left empty
- Duplicate records
- Emails and phone numbers in the wrong format
- Dates and numbers that aren't valid
- Columns that contradict each other
- Broken business rules
- Values missing from reference or master data
Writing separate checks for each dataset means a lot of hard-coded logic, and three problems that grow with every new source.
Inconsistent Validation
Each dataset ends up checked a different way, so "clean" means something different from one team to the next.
Code Changes for Every Rule Change
When the business updates a rule, a developer has to find it, change it and redeploy it.
A Pass or Fail That Explains Nothing
A single result doesn't say which record failed, which rule it broke, or what needs fixing.
One Engine. Your Rules Live in Configuration.
The framework separates what to check from how checks run. Validation rules sit in a configuration file, and a single PySpark engine applies them to any dataset you point it at.
-
Input Data
The dataset you want to validate, loaded into a Spark DataFrame.
-
Configuration
Which rules apply, to which columns, in what order, with what failure message.
-
DQ Engine
The reusable PySpark engine that runs the enabled rules by priority.
-
Results
A status and a plain-language remark added to every record.
-
Audit
Execution details and summary metrics, stored for every run.
A new dataset needs a new configuration, not new validation code.
Watch the Walkthrough
An employee dataset taken from configuration to execution, results and audit.
Seven Rule Types, Ready to Configure
Together they cover the data quality dimensions that matter most: completeness, uniqueness, validity, conformity, consistency, integrity and text quality.
| Rule | What it catches | Dimension |
|---|---|---|
| Null and empty | Missing values in mandatory fields | Completeness |
| Duplicate | Records that appear more than once | Uniqueness |
| Datatype | Values that don't match the expected type, like text in a numeric field | Validity |
| Regex | Badly formatted emails, phone numbers, PAN numbers or postal codes | Conformity |
| Expression | Business rules and relationships between columns | Consistency |
| SQL or lookup | Values that don't exist in reference or master data | Integrity |
| Language | Text and character-level requirements | Text quality |
Rules Are Configured, Not Hard-Coded.
If Employee Name is mandatory, you say so in the configuration file. The engine does the rest. Each rule is described by the same six fields.
- rule_name
- Which check to run
- enabled
- Turn a rule on or off without deleting it
- priority
- The order rules run in
- applies_to
- The columns the rule checks
- parameters
- Settings the rule needs, like a pattern or expression
- fail_message
- The remark written on records that fail
Rules Change. The Engine Doesn't
As business requirements move, you add or edit rules in configuration. The execution flow stays the same.
-
01
Today
Employee Name cannot be null. One null_values_check entry covers it.
-
02
Tomorrow
The business wants Employee Name in a specific format too. Add a format rule to the configuration and run again.
-
03
Later
A check the library doesn't support yet? Add the rule type to the framework once, then use it from configuration on any dataset.
Same Engine, Any Dataset.
The configuration also says which dataset each rule belongs to, so one framework serves every domain.
- Employee
- Recruitment
- Finance
- Sales
- Production
- Insurance
What Happens When You Run It
You don’t run validations one by one. The framework handles the whole lifecycle in a single execution.
-
Load the Data
The input dataset is read into a Spark DataFrame.
-
Read the Configuration
The framework finds the configuration for this dataset.
-
Select Enabled Rules
Only rules switched on for this dataset are picked up.
-
Run by Priority
Rules execute in the order their priority sets.
-
Mark Every Record
Each row gets a validation status and, if it failed, a remark.
-
Record the Run
Execution metrics are calculated and the summary is saved to Delta.
Know Which Record Failed, and Why
Instead of “this dataset has bad data,” every record carries one of three statuses, so data owners know exactly where to look.
Pass
The record passed every configured rule.
Reject
The record failed one or more rules.
Duplicate
The record was identified as a duplicate.
The Remark Tells You What to Fix.
Every rejected record gets a validation remark written from the rule's failure message. That makes the output something a business user can act on, not just a developer.
- R
Invalid email - R
CTC required for Confirmed Employee - R
DOB must be a valid date - R
Experience must be integer
Every Run Leaves a Record.
Beyond record-level results, each execution gets its own ID and summary metrics, persisted to Delta. You can see how quality changes across runs and datasets over time.
Data quality isn't a one-time check. It's something you measure.
- Execution ID
- dq-0924-0831
- Rows processed
- 6
- Rows passed
- 1
- Rules applied
- 5
- Rows flagged
- 5
- Runtime
- 4.8 s
Failures by Rule
- datatype_check2
- null_values_check1
- regex_check1
- duplicate_check1
- expression_check0
Built Once, Used Everywhere
Find bad data early, understand exactly what’s wrong, and give the people who own it what they need to fix it.
Reusable
One engine validates every dataset you configure.
Scalable
Built on PySpark to handle large datasets.
Auditable
Each execution is logged to Delta with its metrics.
Maintainable
Business rules change in configuration, not in code.
Actionable
Every failed record says what went wrong.
Standardized
The same validation approach across teams and projects.
Frequently Asked Questions
How the Data Quality Framework validates data, how rules are managed, and how results are tracked.
A data quality framework is a standard way to check data for errors before it reaches reports, dashboards or business processes. Our Data Quality Framework accelerator is a reusable, configuration-driven engine built on PySpark. It applies validation rules to any dataset and flags exactly which records fail and why.
Instead of writing separate validation code for each dataset, you define rules in a configuration file and let one PySpark engine run them. The framework loads data into a Spark DataFrame, applies the enabled rules by priority, and adds a validation status and remark to every record.
It supports seven rule types: null and empty checks, duplicate checks, datatype validation, regex validation for formats like email, phone, PAN and postal codes, expression rules for business logic, SQL or lookup checks against reference data, and language validation. Together they cover completeness, uniqueness, validity, conformity, consistency, integrity and text quality.
Yes. Rules are configured, not hard-coded. You can add, edit or turn off a rule by updating its name, priority, target columns, parameters and failure message in configuration. If you need a completely new type of check, it's added to the framework once and then reused across any dataset.
Yes. The configuration tells the engine which rules belong to which dataset, so the same framework can validate employee, recruitment, finance, sales, production, insurance or other data. Each new dataset needs its own configuration, not new validation code.
Every run gets a unique execution ID and a summary of rows processed, rows passed, rules applied, failures per rule and runtime. This summary is stored in Delta, building an audit history, so you can see how quality changes across runs and datasets.
Don't Wait for Bad Data to Reach the Business.
See the framework run on a real dataset, from configuration to audit, and talk through how it fits your data platform.