The hardware and bandwidth for this mirror is donated by dogado GmbH, the Webhosting and Full Service-Cloud Provider. Check out our Wordpress Tutorial.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]dogado.de.
A clustering method applied to continuous data will report some
number of clusters whether or not any exist. So when an analysis
announces that a dataset contains, say, four types, the number may
describe the population, or only the shape of the data under that
method. matchednull separates the two.
copula_null() builds a synthetic twin of your data that
keeps everything innocent and removes the one thing in question:
If your pipeline finds as many clusters in the twins as in the real data, those clusters were the data’s shape, not its people.
library(matchednull)
set.seed(1)
x <- matrix(rnorm(400 * 3), 400, 3) %*%
chol(matrix(c(1, .5, .3, .5, 1, .4, .3, .4, 1), 3, 3))
twin <- copula_null(x)
all(sort(twin[, 1]) == sort(x[, 1])) # margins: identical
#> [1] TRUE
round(cor(x) - cor(twin), 2) # correlations: close
#> [,1] [,2] [,3]
#> [1,] 0.00 0.01 0.05
#> [2,] 0.01 0.00 -0.03
#> [3,] 0.05 -0.03 0.00matched_null_test() takes your data and your own
pipeline, wrapped as a function that returns one number
(typically the selected number of clusters). It runs the identical
pipeline on the real data and on R twins, and asks whether
the real answer stands out.
suppressPackageStartupMessages(library(mclust))
pick_k <- function(d) Mclust(d, G = 1:4, modelNames = "VVV", verbose = FALSE)$G
# 1. Typeless data: the test should stay quiet.
set.seed(7)
matched_null_test(x, pick_k, R = 30)
#> Matched-null test (30 null twins)
#> real statistic: 1
#> null interval: [1, 1]
#> p (real >= nulls): 1
#> verdict: null-like (within the twins' interval)# 2. Two genuine types, hidden in the dependence structure
# (identical margins, opposite correlation orientation).
set.seed(42)
z <- sample(2, 400, replace = TRUE)
X <- matrix(rnorm(400 * 4), 400, 4)
L1 <- chol(matrix(c(1, .85, .85, 1), 2, 2))
L2 <- chol(matrix(c(1, -.85, -.85, 1), 2, 2))
X[z == 1, 1:2] <- X[z == 1, 1:2] %*% L1
X[z == 1, 3:4] <- X[z == 1, 3:4] %*% L1
X[z == 2, 1:2] <- X[z == 2, 1:2] %*% L2
X[z == 2, 3:4] <- X[z == 2, 3:4] %*% L2
set.seed(7)
matched_null_test(X, pick_k, R = 30)
#> Matched-null test (30 null twins)
#> real statistic: 2
#> null interval: [1, 1]
#> p (real >= nulls): 0.032
#> verdict: exceeds the null (beyond margins + covariance)The first call returns null-like: whatever clustering the pipeline reports is reproduced by twins with no types. The second returns exceeds the null: the grouping lives in structure the twins cannot carry.
cluster_fn can wrap anything: a k-means heuristic, a
published typology’s exact workflow, or any scalar measure of clustering
strength. The null is defined at the level of the data, not the
pipeline.
Two practical notes:
The method, its positive controls, and its false-positive calibration are described in the accompanying paper: Types Without Taxa: A Covariance-Matched-Null Multiverse Test of Categorical versus Continuous Personality Structure.
These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.
Health stats visible at Monitor.