The hardware and bandwidth for this mirror is donated by dogado GmbH, the Webhosting and Full Service-Cloud Provider. Check out our Wordpress Tutorial.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]dogado.de.

Name-blind variable-role detection with rolescry

The problem: names lie, data does not

Real datasets arrive with column names that are missing, misleading, in the wrong language, or simply wrong: a column called category_code holding continuous lab values, a gender column that is actually a free numeric measurement, an outcome buried under an opaque v7. Any tool that decides a column’s statistical role from its name inherits every one of those lies.

rolescry decides roles from the data signature instead – the guiding principle is Data inspice, non nomen (“inspect the data, not the name”). Renaming every column to col_1, col_2, ... does not change a single role assignment. This is the turnusol (litmus) invariant, and it is the package’s keystone test.

A worked example

set.seed(1)
pre  <- rnorm(100, 10, 2)          # a measurement, before ...
d <- data.frame(
  arm  = rep(c("control", "treatment"), each = 50),  # a balanced 2-level grouping
  pre  = pre,
  post = pre + rnorm(100, 1, 1)                       # ... and again after (paired with pre)
)

res <- detect_roles(d)
res
#> <role_detection> 100 observations x 3 variables
#>   group_var          arm                          pct=100.0
#>   paired_pairs       pre, post                    pct=77.3
#>   agreement_pairs    pre, post                    pct=79.1
#>   covariate          pre, post, arm               pct=50.0
summary(res)
#>                  role found      columns   pct
#> 1           group_var  TRUE          arm 100.0
#> 2        paired_pairs  TRUE     pre,post  77.3
#> 3     agreement_pairs  TRUE     pre,post  79.1
#> 4       time_variable FALSE                0.0
#> 5      event_variable FALSE                0.0
#> 6          subject_id FALSE                0.0
#> 7   repeated_measures FALSE                0.0
#> 8         scale_items FALSE                0.0
#> 9  outcome_continuous FALSE                0.0
#> 10     outcome_binary FALSE                0.0
#> 11          covariate  TRUE pre,post,arm  50.0

The same call on the name-stripped twin yields the same roles by position:

d_blind <- setNames(d, paste0("col_", seq_along(d)))
pos <- function(r, dat) match(r$roles$paired_pairs$columns, names(dat))
identical(pos(detect_roles(d), d), pos(detect_roles(d_blind), d_blind))
#> [1] TRUE

How a role is scored

Each column is first typed by value (continuous, binary, categorical, ID), never by name. Candidate roles are then scored by signatures that capture the statistical shape a role implies – correlation and distributional overlap for paired measurements, Bland-Altman bias and intraclass correlation for agreement, event-rate and right-skew for survival, inter-item correlation and a Cronbach-alpha proxy for scale items, and so on. Every score is a transparent sum of named components you can inspect:

res$roles$paired_pairs$components[[1]]
#> $name
#> [1] "Correlation"
#> 
#> $score
#> [1] 20
#> 
#> $max
#> [1] 20
#> 
#> $detail
#> [1] "r=0.88"

Shannon entropy

For a categorical column with level proportions \(p_1, \dots, p_k\), the normalized Shannon entropy

\[ H_{\text{norm}} = \frac{-\sum_i p_i \log_2 p_i}{\log_2 k} \in [0, 1] \]

measures how balanced the levels are. A grouping variable (treatment vs control) has high entropy (near-balanced); a near-constant flag has entropy near zero. Entropy drives both the value classifier and the group-balance signal.

Normalized mutual information

To ask – name-blind – whether a candidate grouping actually carries information about an outcome, rolescry uses normalized mutual information:

\[ \text{NMI}(X, Y) = \frac{I(X; Y)}{\min\{H(X),\, H(Y)\}} \in [0, 1], \]

which is 0 for independent variables and 1 for a deterministic association, and is comparable across variables with different numbers of levels. It is exposed directly:

set.seed(2)
g <- sample(c("A", "B", "C"), 300, replace = TRUE)
y <- ifelse(g == "A", "event", sample(c("event", "none"), 300, replace = TRUE))
compute_nmi(g, y)          # > 0: g informs y
#> [1] 0.2893517
compute_nmi(g, sample(g))  # ~ 0: shuffled -> independent
#> [1] 0.001503439

The optional, capped name bonus

Names are not useless – they are just untrustworthy. When you do trust them, pass a keyword dictionary via name_bonus. Names then act only as a small, capped tie-breaker (at most a +10 point nudge, i.e. <= 10% of the selection score); the mathematical signature still dominates (>= 90%), the relationship enforced by score_gap_ok().

Here two columns both look like groupings; the data alone picks the perfectly balanced site, but a keyword dictionary nudges the choice to the intended treatment arm:

set.seed(4)
clin <- data.frame(
  site    = rep(c("north", "south"), length.out = 160),  # perfectly balanced (the math pick)
  treated = sample(c("no", "yes"), 160, replace = TRUE, prob = c(0.62, 0.38))
)
detect_roles(clin)$roles$group_var$columns                                  # data alone -> "site"
#> [1] "site"
detect_roles(clin, name_bonus = rolescry_default_name_bonus())$roles$group_var$columns  # name nudge -> "treated"
#> [1] "treated"

Header-aware loading

read_data() reads a file with the header row found by the same information-theoretic scorer (detect_header()), so messy exports with title rows or merged cells still load with sensible column names. Delimited text works with base R; spreadsheet and statistical formats use optional packages and degrade gracefully if they are not installed.

Attribution

rolescry is derived from Boynukara, C. (2026). MDStatR (v2.1.0 Veritas). Zenodo. https://doi.org/10.5281/zenodo.20707791. Run citation("rolescry") to cite the package and its parent engine.

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.
Health stats visible at Monitor.