The hardware and bandwidth for this mirror is donated by dogado GmbH, the Webhosting and Full Service-Cloud Provider. Check out our Wordpress Tutorial.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]dogado.de.

Nested and longitudinal data

Long format, natively

Repeated-measures and nested designs are the reason flexsynth works in long format. You never pivot to wide: the structure formula declares the hierarchy, for example ~ id / visit (patients, each with a sequence of visits).

We simulate a longitudinal cohort. Each patient has a fixed baseline (age, sex) and a time-varying systolic blood pressure that drifts over visits — an autocorrelated series with an unequal number of visits per patient.

n_pat <- 120
base <- data.frame(
  id  = seq_len(n_pat),
  age = round(rnorm(n_pat, 60, 10)),
  sex = sample(c("F", "M"), n_pat, replace = TRUE)
)
n_visits <- 1 + rpois(n_pat, 2)                 # unequal length units

rows <- lapply(seq_len(n_pat), function(i) {
  k <- n_visits[i]
  sbp <- numeric(k)
  sbp[1] <- 0.6 * base$age[i] + rnorm(1, 95, 8)
  for (t in seq_len(k)[-1]) sbp[t] <- 0.8 * sbp[t - 1] + rnorm(1, 26, 6)
  data.frame(id = base$id[i], visit = seq_len(k),
             age = base$age[i], sex = base$sex[i],
             sbp = round(sbp))
})
real <- do.call(rbind, rows)
head(real, 8)
id visit age sex sbp
1 1 54 M 133
2 1 62 M 140
2 2 62 M 140
3 1 52 F 119
3 2 52 F 128
3 3 52 F 117
3 4 52 F 116
4 1 76 M 139

Synthesising the hierarchy

res <- synth(real, structure = ~ id / visit, seed = 1)
res
#> <synth_result>
#>   track        : A (high-utility; NOT differentially private)
#>   datasets (m) : 1 
#>   rows each    : 358  (input: 366 )
#>   synthesised  : age, sex, sbp 
#>   unit-level   : age, sex (once per unit)
#>   carried      : visit 
#>   method       : cart 
#> 
#> Get the data with as.data.frame(x)  

The engine detects that age and sex are constant within a patient (subject-invariant) and synthesises them once per unit, broadcasting them across that unit’s visits, so a synthetic patient stays internally consistent. The number of visits per patient is drawn from the learned count distribution, and sbp is synthesised with an initial-state model plus a Markov transition that conditions on the previous visit — so the within-patient autocorrelation survives.

syn <- as.data.frame(res)

# unequal-length units are reproduced, not fixed-width
table(table(syn$id))
#> 
#>  1  2  3  4  5  6  7 
#> 19 34 29 14 19  2  3 
# baseline stays constant within each synthetic patient
all(tapply(syn$age, syn$id, function(a) length(unique(a)) == 1))
#> [1] TRUE

Within-unit temporal constraints

rule() with scope = "unit" evaluates a rule per unit, which is how you express temporal logic. Suppose visits must be numbered consecutively from 1 and sbp must stay in a plausible range on every visit:

res_c <- synth(
  real, ~ id / visit,
  constraints = list(
    rule(sbp >= 60 & sbp <= 260),                       # row scope (default)
    rule(all(diff(visit) == 1), scope = "unit")         # consecutive visits
  ),
  seed = 1
)
syn_c <- as.data.frame(res_c)

# every synthetic unit's visits are 1, 2, 3, ...
all(tapply(syn_c$visit, syn_c$id,
           function(v) identical(v, seq_along(v))))
#> [1] TRUE

Units that violate any rule are regenerated (bounded by constraint_max_tries in synth_control()), so the nested structure is never broken to satisfy a constraint.

Diagnostics

The utility and risk diagnostics work unchanged on long data.

diagnose(real, res, vars = c("age", "sex", "sbp"))
#> <flexsynth_diagnostics>
#>   rows        : real 366  synthetic 358 
#>   variables   : 3 
#> 
#> Univariate fit (smaller = closer):
#>  variable        type metric distance
#>       age     numeric     ks   0.1345
#>       sex categorical    tvd   0.0082
#>       sbp     numeric     ks   0.0924
#>   mean distance: 0.0784   worst: age (0.1345)
#> 
#> Correlation structure (2 numeric vars):
#>   Frobenius diff: 0.0555   mean |diff|: 0.0392   max |diff|: 0.0392
#> 
#> Propensity utility (pMSE, logistic; descriptive, in-sample):
#>   pMSE: 0.00035   expected: 0.00052   ratio: 0.68 (1 = indistinguishable)

Because visits share a patient, treat the diagnostics as descriptive: the pMSE here mixes baseline and time-varying columns rather than modelling the dependence explicitly.

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.
Health stats visible at Monitor.