The hardware and bandwidth for this mirror is donated by dogado GmbH, the Webhosting and Full Service-Cloud Provider. Check out our Wordpress Tutorial.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]dogado.de.
An R package for defining causal data generating processes with known ground truth and evaluating estimator performance against them.
"low",
"moderate", "high") or custom functionsRequires R 4.0 or higher. Install from GitHub:
# install.packages("devtools")
devtools::install_github("chaycereed/causalsim")causalsim() generates a dataset in one call. The
returned data frame includes covariate columns, treatment
A, outcome Y, individual effect
.tau, and true propensity .p.
library(causalsim)
data <- causalsim(n = 500, n_confounders = 1, effect = 2, seed = 1L)
head(data)For estimator benchmarking or multi-draw workflows, use
causalsim_dgp() and causalsim_draw() directly
(see below).
causalsim_dgp() specifies the structural model.
Parameters can be scalars, preset strings, or functions of the covariate
names.
dgp <- causalsim_dgp(
n = 500,
n_confounders = 1,
effect = 2,
propensity = "moderate",
baseline = "moderate"
)
dgpcausalsim_draw() simulates one dataset from the DGP. The
returned data frame includes covariate columns, treatment
A, outcome Y, individual effect
.tau, and propensity .p.
dat <- causalsim_draw(dgp, seed = 1L)
head(dat)An estimator is any function that accepts a data frame and returns a
named numeric vector with at minimum an estimate field.
ci_lower and ci_upper enable coverage and
power metrics.
ols_est <- function(data) {
fit <- lm(Y ~ A + W, data = data)
est <- coef(fit)[["A"]]
se <- sqrt(vcov(fit)["A", "A"])
c(estimate = est, ci_lower = est - 1.96 * se, ci_upper = est + 1.96 * se)
}
result <- causalsim_eval(dgp, ols_est, reps = 200L, seed = 1L)
result
summary(result)
plot(result)causalsim_eval_grid() runs the evaluator over the
Cartesian product of any DGP parameters, returning a tidy data frame of
metrics for each cell.
grid_result <- causalsim_eval_grid(
dgp = dgp,
estimator = ols_est,
vary = list(n = c(100L, 250L, 500L, 1000L)),
reps = 200L,
metrics = c("bias", "rmse"),
seed = 1L
)
grid_resultDeclare a covariate with role = "effect_modifier" and
reference it in a function passed to effect to make the
treatment effect vary across subgroups. The individual effects are
stored in the .tau column as ground truth.
het_dgp <- causalsim_dgp(
n = 2000,
covariates = list(
W = causalsim_covar("normal", role = "confounder"),
V = causalsim_covar("binary", role = "effect_modifier", prob = 0.5)
),
effect = function(V) 2 + 3 * V # effect is 2 when V = 0, 5 when V = 1
)
d <- causalsim_draw(het_dgp, seed = 1L)
tapply(d$.tau, d$V, mean) # 2 for V = 0, 5 for V = 1The function passed to effect is what activates the
modifier; a covariate labelled effect_modifier that
effect never references is inert, and
causalsim_dgp() warns when that happens.
Planned for a future release:
confounder, effect_modifier,
noise) cover the common cases. Two single-path roles would
complete the treatment/outcome taxonomy:
instrument — drives treatment only (enters the
propensity model, excluded from the outcome), for benchmarking
instrumental-variable estimators.prognostic — drives the outcome only (enters the
baseline, independent of treatment), for studying precision covariates
and variance reduction.MIT License. See LICENSE for details.
These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.
Health stats visible at Monitor.