The hardware and bandwidth for this mirror is donated by dogado GmbH, the Webhosting and Full Service-Cloud Provider. Check out our Wordpress Tutorial.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]dogado.de.
Data quality is an important prerequisite for reliable statistical analysis, visualization, modelling, and interpretation.
DataAudit provides a structured framework for
identifying common data-quality problems in R. The package combines
general data auditing, individual validation functions, reusable
validation rules, structured audit reports, and data-quality
scoring.
A typical workflow is:
Consider a small dataset containing several intentional data-quality problems.
dat <- data.frame(
ID = c(
"A001",
"A002",
"A002",
"BAD",
"A005"
),
Species = c(
"Dog",
"Horse",
"",
"Cat",
"Dog"
),
Age = c(
5,
40,
-2,
10,
NA
),
stringsAsFactors = FALSE
)
dat
#> ID Species Age
#> 1 A001 Dog 5
#> 2 A002 Horse 40
#> 3 A002 -2
#> 4 BAD Cat 10
#> 5 A005 Dog NAThe dataset contains examples of:
The primary high-level function is audit_data().
report <- audit_data(
dat,
include_nzv = FALSE
)
report
#> DataAudit Report
#> ================
#>
#> Rows: 5
#> Variables: 3
#> Variables with issues: 2
#> Duplicated rows: 0
#> Total issues: 2
#> Rule violations: 0
#> Overall issues: 2
#>
#> Variables with issues
#> ---------------------
#> Variable Class Missing MissingPercent Blank Infinite Unique Constant
#> Species character 0 0 1 0 4 FALSE
#> Age numeric 1 20 0 0 4 FALSE
#> NearZeroVariance Issue
#> FALSE TRUE
#> FALSE TRUE
#>
#> No duplicated rows detected.The resulting object inherits from the DataAuditReport
class.
The report contains several components.
These provide information about:
The overview component provides high-level audit statistics.
Variable-level diagnostics are stored in the variables
component.
report$variables
#> Variable Class Missing MissingPercent Blank Infinite Unique
#> ID ID character 0 0 0 0 4
#> Species Species character 0 0 1 0 4
#> Age Age numeric 1 20 0 0 4
#> Constant NearZeroVariance Issue
#> ID FALSE FALSE FALSE
#> Species FALSE FALSE TRUE
#> Age FALSE FALSE TRUEThis provides information such as missing values, blank values, infinite values, numbers of unique values, constant variables, near-zero variance variables, and whether a variable contains a detected issue.
Duplicated observations can be inspected through:
This checks duplicated rows. Identifier uniqueness can be assessed
separately using unique_check() or through validation
rules.
DataAudit functions can also be used independently.
Suppose valid ages are expected to range from 0 to 30.
Suppose only "Dog" and "Cat" are accepted
categories.
Suppose identifiers must begin with A followed by three
digits.
For repeated or study-specific validation, multiple requirements can
be defined together using audit_rules().
rules <- audit_rules(
range = list(
Age = c(0, 30)
),
category = list(
Species = c(
"Dog",
"Cat"
)
),
unique = "ID",
required = "Species",
pattern = list(
ID = "^A[0-9]{3}$"
)
)
rules
#> $range
#> $range$Age
#> [1] 0 30
#>
#>
#> $category
#> $category$Species
#> [1] "Dog" "Cat"
#>
#>
#> $unique
#> [1] "ID"
#>
#> $required
#> [1] "Species"
#>
#> $pattern
#> $pattern$ID
#> [1] "^A[0-9]{3}$"
#>
#>
#> attr(,"class")
#> [1] "DataAuditRules" "list"The rules object can be reused with datasets having the same expected structure.
Rules can be evaluated directly using
apply_audit_rules().
rule_results <- apply_audit_rules(
dat,
rules
)
rule_results
#> DataAudit Rule Validation
#> =========================
#>
#> Status: Violations detected
#> Total violations: 8
#>
#> Rule summary
#> ------------
#> Rule Violations
#> Range 2
#> Category 2
#> Unique 2
#> Required 1
#> Pattern 1
#>
#> Variables with rule violations
#> ------------------------------
#> - Range: Age
#> - Category: Species
#> - Unique: ID
#> - Required: Species
#> - Pattern: IDThe returned object contains the rule-specific results and total number of violations.
rule_results$summary
#> Rule Violations
#> 1 Range 2
#> 2 Category 2
#> 3 Unique 2
#> 4 Required 1
#> 5 Pattern 1
rule_results$total_violations
#> [1] 8A compact summary can also be obtained using:
summary(rule_results)
#> $status
#> [1] "Violations detected"
#>
#> $total_violations
#> [1] 8
#>
#> $by_rule
#> Rule Violations
#> 1 Range 2
#> 2 Category 2
#> 3 Unique 2
#> 4 Required 1
#> 5 Pattern 1
#>
#> $violations
#> Rule Variable Row Value
#> 1 Range Age 2 40
#> 2 Range Age 3 -2
#> 3 Category Species 2 Horse
#> 4 Category Species 3
#> 5 Unique ID 2 A002
#> 6 Unique ID 3 A002
#> 7 Required Species 3
#> 8 Pattern ID 4 BAD
#>
#> attr(,"class")
#> [1] "summary.DataAuditRuleResult" "list"The validation rules can be incorporated directly into
audit_data().
report <- audit_data(
dat,
include_nzv = FALSE,
rules = rules
)
report
#> DataAudit Report
#> ================
#>
#> Rows: 5
#> Variables: 3
#> Variables with issues: 2
#> Duplicated rows: 0
#> Total issues: 2
#> Rule violations: 8
#> Overall issues: 10
#>
#> Variables with issues
#> ---------------------
#> Variable Class Missing MissingPercent Blank Infinite Unique Constant
#> Species character 0 0 1 0 4 FALSE
#> Age numeric 1 20 0 0 4 FALSE
#> NearZeroVariance Issue
#> FALSE TRUE
#> FALSE TRUE
#>
#> No duplicated rows detected.
#>
#> Custom rule validation
#> ----------------------
#> Rule Violations
#> Range 2
#> Category 2
#> Unique 2
#> Required 1
#> Pattern 1
#>
#> Total rule violations: 8The resulting report now combines automatic data-quality assessment with user-defined validation rules.
report$summary
#> Rows Variables VariablesWithIssues DuplicateRows TotalIssues RuleViolations
#> 1 5 3 2 0 2 8
#> OverallIssues
#> 1 10
report$rules
#> DataAudit Rule Validation
#> =========================
#>
#> Status: Violations detected
#> Total violations: 8
#>
#> Rule summary
#> ------------
#> Rule Violations
#> Range 2
#> Category 2
#> Unique 2
#> Required 1
#> Pattern 1
#>
#> Variables with rule violations
#> ------------------------------
#> - Range: Age
#> - Category: Species
#> - Unique: ID
#> - Required: Species
#> - Pattern: IDThe report can be summarized using the standard
summary() generic.
audit_summary <- summary(report)
audit_summary
#> $dataset
#> Rows Variables VariablesWithIssues DuplicateRows TotalIssues RuleViolations
#> 1 5 3 2 0 2 8
#> OverallIssues
#> 1 10
#>
#> $variables_with_issues
#> Variable Class Missing MissingPercent Blank Infinite Unique
#> Species Species character 0 0 1 0 4
#> Age Age numeric 1 20 0 0 4
#> Constant NearZeroVariance Issue
#> Species FALSE FALSE TRUE
#> Age FALSE FALSE TRUE
#>
#> $duplicate_rows
#> [1] Row Duplicate
#> <0 rows> (or 0-length row.names)
#>
#> $rules
#> Rule Violations
#> 1 Range 2
#> 2 Category 2
#> 3 Unique 2
#> 4 Required 1
#> 5 Pattern 1
#>
#> attr(,"class")
#> [1] "summary.DataAuditReport" "list"The summary object provides compact access to dataset-level information, variables with issues, duplicated rows, and rule results.
audit_summary$dataset
#> Rows Variables VariablesWithIssues DuplicateRows TotalIssues RuleViolations
#> 1 5 3 2 0 2 8
#> OverallIssues
#> 1 10
audit_summary$variables_with_issues
#> Variable Class Missing MissingPercent Blank Infinite Unique
#> Species Species character 0 0 1 0 4
#> Age Age numeric 1 20 0 0 4
#> Constant NearZeroVariance Issue
#> Species FALSE FALSE TRUE
#> Age FALSE FALSE TRUE
audit_summary$duplicate_rows
#> [1] Row Duplicate
#> <0 rows> (or 0-length row.names)
audit_summary$rules
#> Rule Violations
#> 1 Range 2
#> 2 Category 2
#> 3 Unique 2
#> 4 Required 1
#> 5 Pattern 1An overall data-quality score can be calculated from a
DataAuditReport.
score <- audit_score(report)
score
#> DataAudit Quality Score
#> =======================
#>
#> Score: 60.00 / 100
#> Quality: Poor
#>
#> Dataset
#> -------
#> Rows: 5
#> Variables: 3
#> Data cells: 15
#>
#> Issues
#> ------
#> Automatic issues: 2
#> Rule violations: 8
#> Overall issues: 10The score ranges from 0 to 100, with higher values indicating fewer detected issues relative to dataset size.
The score is accompanied by one of five qualitative categories:
The individual components can be inspected directly.
score$score
#> [1] 60
score$quality
#> [1] "Poor"
score$automatic_issues
#> [1] 2
score$rule_violations
#> [1] 8
score$overall_issues
#> [1] 10The data-quality score is intended as a compact audit summary. It should not replace domain-specific assessment of whether a dataset is suitable for a particular scientific or analytical purpose.
DataAudit provides additional focused functions for different aspects of data quality, including:
constant_check();nzv_check();infinite_check();outlier_check();type_check();date_check();id_check();length_check();case_check();whitespace_check();sequence_check();group_sequence_check();dependency_check(); andconsistency_check().These functions can be combined with the general audit when more detailed validation is required.
A practical workflow is:
Raw data
|
v
audit_data()
|
+---- General data-quality assessment
|
+---- Individual checks
|
v
audit_rules()
|
v
apply_audit_rules()
|
v
Integrated DataAuditReport
|
+---- print()
|
+---- summary()
|
v
audit_score()
This separates general data-quality screening from domain-specific validation while allowing both to be incorporated into a single structured report.
DataAudit provides a reproducible framework for
detecting, organizing, and summarizing common data-quality problems in
R.
The package can be used for rapid preliminary screening with
audit_data(), focused validation through individual check
functions, reusable validation with audit_rules(), and
compact reporting through summary() and
audit_score().
These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.
Health stats visible at Monitor.