The hardware and bandwidth for this mirror is donated by dogado GmbH, the Webhosting and Full Service-Cloud Provider. Check out our Wordpress Tutorial.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]dogado.de.

R build status CRAN_Status_Badge Downloads Total Downloads

datarium: Data bank for statistical analyses and visualization datarium hex logo

Contains 23 curated data sets for teaching classical statistical inference in R. Each data set is paired with the analysis it is designed to teach — one data set per test. The same map is available in R via ?datarium. Each analysis below links to its worked tutorial on Datanovia.

Comparing means

Analysis Data set
One-sample t-test mice
Paired-samples t-test mice2
Two-samples (independent) t-test genderweight

Analysis of variance (ANOVA)

Analysis Data set
Two-way ANOVA jobsatisfaction
Three-way ANOVA headache, heartattack
One-way repeated measures ANOVA selfesteem
Two-way repeated measures ANOVA selfesteem2
Three-way repeated measures ANOVA weightloss
Two-way mixed ANOVA anxiety, depression
Three-way mixed ANOVA performance
Two-way ANCOVA stress

Categorical data and proportions

Analysis Data set
Chi-square test of independence properties, housetasks.raw
Categorical descriptive statistics titanic.raw
McNemar’s test antismoking
Cochran’s Q test taskachievment
Cochran-Armitage trend test renalstone

Regression

Analysis Data set
Linear regression marketing
Poisson regression (count data) infections
Logistic regression (real, messy clinical data) heartdisease

Time series

Analysis Data set
Time series AirPassengersDf

Helper functions

The two data sets (Titanic and housetasks) are frequency/contingency table. We’ll create our demo data sets by recovering the original data from Titanic and housetasks tables.

To do so, first copy and paste the following helper function:

counts_to_cases <- function(x, countcol = "Freq") {
    if(!inherits(x, "table")) x <- as.table(as.matrix(x))
      x <- as.data.frame(x)
    # Get the row indices to pull from x
    idx <- rep.int(seq_len(nrow(x)), x[[countcol]])
    # Drop count column
    x[[countcol]] <- NULL
    # Get the rows from x
    x <- x[idx, ]
    rownames(x) <- 1:nrow(x)
    x
}

Then, recover the original data as follow:

# Load the data
data("Titanic")
data("housetasks", package = "factoextra")

# Recover the original raw data
titanic.raw <- counts_to_cases(Titanic)
housetasks.raw <- counts_to_cases(housetasks)

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.
Health stats visible at Monitor.