The hardware and bandwidth for this mirror is donated by dogado GmbH, the Webhosting and Full Service-Cloud Provider. Check out our Wordpress Tutorial.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]dogado.de.

Was silence informative? The silence test, the sensor-gap test and the fatigue check

Hsiu-Ting Yu

The question

A declared dm-graph says which mechanism is assumed (see vignette("dm-graphs")). The tests in this vignette ask the data whether that assumption is tenable. They exploit one fact: under every recoverable mechanism, the response indicator at prompt t carries no information about the states beyond what the observed history already carries. If skipped prompts turn out to predict what is observed afterwards, silence was informative.

Two tests and one descriptive check are available.

Function Regression (within person) Null holds under Needs
silence_test() X_{t+1} ~ X_{t-1} + R_t over triples with t-1 and t+1 answered M0, M1, M3, M4 and combinations without M1 + M3 nothing beyond the states
sensor_gap_test() S_t ~ X_{t-1} + R_t over prompts with an answered predecessor all recoverable motifs, including M1 + M3 an always-observed sensor
fatigue_check() response rate after an answered vs a skipped prompt; by study quarter (descriptive) the response indicators

The two tests use person fixed effects (within-person demeaning), so stable between-person differences in level or in compliance cannot produce a spurious result. The fatigue check reports both a pooled contrast, which between-person differences in compliance inflate, and the mean of the person-specific contrasts, which they do not.

The silence test

Take every triple of consecutive scheduled prompts (t-1, t, t+1) in which the two outer prompts were answered. The middle prompt may or may not have been answered. Regress the state at t+1 on the state at t-1 and the response indicator at t, within person. If the middle skip is uninformative, its coefficient is zero.

sim2 <- simulate_ema(N = 60, n_prompts = 40, motifs = "M2", compliance = 0.7, delta = -1, seed = 1)
st <- silence_test(sim2$data, sim2$vars)
st
#> Silence test: X_{t+1} ~ X_{t-1} + R_t (within person), 1187 triples, 226 with a skipped middle prompt
#>  variable coef_R    se     z       p
#>      NegA -0.299 0.075 -4.00 0.00018
#>      PosA  0.161 0.076  2.13 0.03800
#>    Stress -0.147 0.070 -2.12 0.03800
#>   Fatigue  0.030 0.077  0.39 0.70000
#> Joint Wald test: statistic = 17.84 on 4 df, p = 0.00324   [SE: cluster; t / F with G - 1 df reference]
#> Interpretation: coefficients near zero are expected under the recoverable motifs; a negative coefficient on a variable means the states hidden by skips were higher.

The data were generated with self-censoring on negative affect (delta = -1: the probit propensity to respond falls by one unit per unit of negative affect). The coefficient of R_t for NegA is negative: after a skipped prompt (R_t = 0), negative affect at t+1 is higher than after an answered one, given the state at t-1. The states hidden by skips were higher than the reported ones, which is exactly what self-censoring produces. The joint Wald test over all four variables combines the evidence.

The result object is a data frame with attributes:

attr(st, "joint")
#>        chisq           df            p 
#> 17.841240732  4.000000000  0.003235468
c(triples = attr(st, "n_triples"), skipped_middle = attr(st, "n_skipped"))
#>        triples skipped_middle 
#>           1187            226

Under a recoverable mechanism the coefficients are near zero and the joint test does not reject:

sim1 <- simulate_ema(N = 60, n_prompts = 40, motifs = "M1", compliance = 0.7, seed = 1)
silence_test(sim1$data, sim1$vars)
#> Silence test: X_{t+1} ~ X_{t-1} + R_t (within person), 1147 triples, 276 with a skipped middle prompt
#>  variable coef_R    se     z     p
#>      NegA -0.038 0.072 -0.53 0.600
#>      PosA  0.138 0.070  1.96 0.054
#>    Stress  0.035 0.077  0.46 0.650
#>   Fatigue -0.063 0.081 -0.77 0.440
#> Joint Wald test: statistic = 4.82 on 4 df, p = 0.318   [SE: cluster; t / F with G - 1 df reference]
#> Interpretation: coefficients near zero are expected under the recoverable motifs; a negative coefficient on a variable means the states hidden by skips were higher.

Reading the sign

The sign of the coefficient of R_t on the self-censoring variable is the sign of the tilt: a negative coefficient means that the hidden states were higher (people skip when the state is high), a positive one that they were lower. This sign is what tilt_profile() and calibrate_delta() later quantify. Coefficients on the other variables reflect the same tilt propagated through the contemporaneous and lagged associations.

Inference

With at least 10 persons the default is a cluster-robust covariance by person with t(G - 1) and F(q, G - 1) reference distributions (Cameron & Miller, 2015). For a single person (or very few), se = "dayblock" resamples person-day blocks with replacement; it needs a day column.

one <- simulate_ema(N = 1, n_prompts = 200, motifs = "M2", compliance = 0.7, delta = -1, days = 8, seed = 5)
silence_test(one$data, one$vars, day = "day", B = 100, seed = 1)
#> Silence test: X_{t+1} ~ X_{t-1} + R_t (within person), 73 triples, 15 with a skipped middle prompt
#>  variable coef_R    se     z    p
#>      NegA -0.359 0.314 -1.14 0.25
#>      PosA  0.194 0.216  0.90 0.37
#>    Stress  0.145 0.346  0.42 0.68
#>   Fatigue -0.018 0.327 -0.05 0.96
#> Joint Wald test: statistic = 3.57 on 4 df, p = 0.467   [SE: dayblock; normal / chi-square reference]
#> Interpretation: coefficients near zero are expected under the recoverable motifs; a negative coefficient on a variable means the states hidden by skips were higher.

With one person and a few dozen triples the test has little power; the coefficient for NegA has the expected sign but its bootstrap interval covers zero. Single-case analyses should lean on the sensitivity analysis rather than on the test.

Curvature

Under M1 the selection on X_t through R_{t+1} makes the conditional mean of X_{t+1} given X_{t-1} slightly nonlinear; poly = 2 adds squared lagged states. In the simulation studies of Yu (2026) the linear test kept its size within Monte Carlo error under M1, so poly = 1 is the default.

silence_test(sim1$data, sim1$vars, poly = 2)$p
#> [1] 0.63091580 0.05951182 0.61709757 0.41464900

The M1 + M3 collider

When lagged-state dependence and burden are both present, conditioning on an answered prompt at t+1 opens a collider at R_{t+1}: its parents are R_t (burden) and X_t (lagged-state dependence), and X_t drives X_{t+1}. Given that the prompt at t+1 was answered despite a skip at t, the state at t must have favored responding, so the state at t+1 is shifted, and the silence test rejects in large samples although the transition kernel is recoverable. recoverability() warns about it. The coefficient has the sign opposite to the self-censoring signature (positive when high states lower the propensity to respond), which is one reason to run fatigue_check() before interpreting a rejection. Under M2 + M3 the analogous collider (with X_{t+1} as a parent of R_{t+1}) works against the self-censoring signal, so the test loses power. In both cases the sensor-gap test is the better instrument.

sim13 <- simulate_ema(N = 300, n_prompts = 56, motifs = c("M1", "M3"), compliance = 0.7, kappa_R = 1, seed = 2)
silence_test(sim13$data, sim13$vars)
#> Silence test: X_{t+1} ~ X_{t-1} + R_t (within person), 8713 triples, 830 with a skipped middle prompt
#>  variable coef_R    se     z      p
#>      NegA  0.122 0.038  3.21 0.0015
#>      PosA -0.073 0.038 -1.94 0.0540
#>    Stress  0.029 0.035  0.82 0.4100
#>   Fatigue  0.011 0.040  0.28 0.7800
#> Joint Wald test: statistic = 12.37 on 4 df, p = 0.0162   [SE: cluster; t / F with G - 1 df reference]
#> Interpretation: coefficients near zero are expected under the recoverable motifs; a negative coefficient on a variable means the states hidden by skips were higher.
recoverability(dm_graph(c("M1", "M3")))$estimator[6]
#> [1] "null violated although the kernel is recoverable (collider at R_{t+1}); use the sensor-gap test or check for fatigue first"

The sensor-gap test

An always-observed channel that loads on the states (an accelerometer, heart rate, phone-usage features) turns the question around: instead of asking what happens after a skip, ask what the sensor recorded at the skipped prompt. sensor_gap_test() regresses the sensor at t on the state at t-1 and R_t, within person, over all prompts whose predecessor was answered. Nothing at t+1 is conditioned on, so the test is valid under M1 + M3 as well.

sim2s <- simulate_ema(N = 60, n_prompts = 40, motifs = "M2", compliance = 0.7, delta = -1,
                      sensor_cor = 0.6, seed = 3)
sg <- sensor_gap_test(sim2s$data, sim2s$vars, sensor = "S")
sg
#> Sensor-gap test: S_t ~ X_{t-1} + R_t (within person), 1640 prompts with answered predecessor
#>   coefficient of R_t: -0.611 (SE 0.052), z = -11.86, p = 2.92e-17   [SE: cluster]
#>   sensor loading on states (answered prompts): 0.571 -0.003 0.025 -0.011
attr(sg, "loading")
#>        NegA        PosA      Stress     Fatigue 
#>  0.57076631 -0.00251733  0.02456381 -0.01115139

The sensor gap is negative: at skipped prompts the sensor recorded higher values of what it measures. The loading attribute gives the within-person regression of the sensor on the states at answered prompts; the calibration in vignette("sensitivity-analysis") uses it to translate the gap into a sensitivity value.

Under a recoverable mechanism the gap is null:

sim1s <- simulate_ema(N = 60, n_prompts = 40, motifs = c("M1", "M3"), compliance = 0.7, sensor_cor = 0.6, seed = 3)
sensor_gap_test(sim1s$data, sim1s$vars, sensor = "S")
#> Sensor-gap test: S_t ~ X_{t-1} + R_t (within person), 1635 prompts with answered predecessor
#>   coefficient of R_t: -0.056 (SE 0.070), z = -0.80, p = 0.425   [SE: cluster]
#>   sensor loading on states (answered prompts): 0.585 0.001 0.029 0.025

The fatigue check

fatigue_check() is descriptive. It reports the response rate after an answered and after a skipped prompt, pooled and as the mean of the person-specific differences, and the response rate by study quarter. A large persistence contrast (the kappa_R term of the simulator) and, when compliance drifts over the study (a negative burden term), a declining rate over quarters are what burden (M3) produces; the example below uses kappa_R only, so the quarterly rates stay flat. Some serial dependence in R also arises under M1, M2 and M4 through the autocorrelated states and the person propensities, so the check does not discriminate M3 from the other motifs by itself.

sim3 <- simulate_ema(N = 60, n_prompts = 40, motifs = "M3", compliance = 0.7, kappa_R = 1, seed = 4)
fatigue_check(sim3$data)
#> Fatigue check over 2340 consecutive prompts
#>   response rate after an answered prompt: 0.808; after a skipped prompt: 0.450; difference 0.358 (person-mean difference 0.327)
#>   response rate by study quarter: 0.712 0.687 0.69 0.71
fatigue_check(sim2$data)
#> Fatigue check over 2340 consecutive prompts
#>   response rate after an answered prompt: 0.770; after a skipped prompt: 0.535; difference 0.235 (person-mean difference 0.111)
#>   response rate by study quarter: 0.702 0.692 0.707 0.7

Under burden (first output) the response rate after a skipped prompt is far below the rate after an answered one, pooled and within person. Under self-censoring alone (second output) the pooled contrast is smaller and the person-mean contrast smaller still: skips cluster in time because the states that cause them are autocorrelated, not because a skip itself lowers the next propensity.

Where burden is declared together with self-censoring, the observed persistence is the quantity that calibrate_delta(method = "postskip", burden = "fit") matches.

How often do the tests reject? A small Monte Carlo

The loop below repeats the silence test on new data sets under three mechanisms. It is deliberately small (the accompanying article uses 500 replications per cell and larger samples); the point is to show the pattern: size near the nominal level under M0 and M1, power under M2.

rej <- function(motifs, reps = 60, ...) {
  mean(vapply(seq_len(reps), function(r) {
    s <- simulate_ema(N = 60, n_prompts = 40, motifs = motifs, compliance = 0.7, seed = 1000 + r, ...)
    attr(silence_test(s$data, s$vars), "joint")["p"] < 0.05
  }, logical(1)))
}
c(M0 = rej("M0"), M1 = rej("M1"), M2 = rej("M2", delta = -1))
#>         M0         M1         M2 
#> 0.08333333 0.08333333 1.00000000

With 60 replications the Monte Carlo standard error of a rejection rate near .05 is about .03, so the first two numbers are compatible with the nominal level; the third shows the power the test has at this sample size and tilt. In the simulation studies of Yu (2026) (500 replications per cell, N = 100, T = 56) the rejection rate under the recoverable motifs stayed between 4% and 7% with two mild excesses (8.8% under burden at 70% compliance and 7.4% under lagged-state dependence at 85%), compatible with Monte Carlo error and the known anticonservatism of cluster-robust standard errors.

What to do with a rejection

A rejection says that the declared recoverable class is not tenable (or, under M1 + M3, that the collider is at work). It does not say which non-recoverable mechanism is present. The recommended continuation is

  1. check for burden with fatigue_check(); if it is substantial and M1 is plausible, prefer the sensor-gap test;
  2. if the sensor-gap test also rejects, or no sensor exists and burden is small, treat the skips as self-censoring and run the sensitivity analysis of vignette("sensitivity-analysis"), using the sign of the silence test as the sign of the tilt;
  3. record both tests in the missingness_declaration().

A non-rejection is weaker evidence: at typical sample sizes the tests have limited power against small tilts, and the identified sets of the sensitivity analysis remain the honest summary when self-censoring is substantively plausible.

References

Cameron, A. C., & Miller, D. L. (2015). A practitioner’s guide to cluster-robust inference. Journal of Human Resources, 50, 317-372. https://doi.org/10.3368/jhr.50.2.317

Yu, H.-T. (2026). What skipped prompts hide: Detecting, diagnosing, and correcting informative nonresponse in ecological momentary assessment. Manuscript under review.

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.
Health stats visible at Monitor.