Skip to contents

About this article: How MockData generates survival dates and why it works this way: where survival dates sit in the generation pipeline, how events and censoring are decided, which layer of the output a column belongs to, and what version 0.5 does not yet do. Each mechanism is shown with code that runs when the vignette is built, and hidden checks stop the build if MockData stops behaving as described. The full design record is the survival-dates decision record, development/adr/v05-survival-dates.md in the MockData repository.

A survival date is derived from its anchor

A survival date depends on another generated column, its anchor. MockData therefore does not draw it with the other variables. Generation runs in four stages:

  1. Baseline: every sampled variable, including the anchor dates.
  2. Survival: each survival date, computed from its anchor.
  3. Formulas: each mockFormula variable, which may use survival dates.
  4. Post-processing: missing codes and garbage, applied to every column.

create_mock_data() runs all four. The stages can also be run one at a time:

spec <- mock_spec(
  mock_spec_date("entry", range = as.Date(c("2001-01-01", "2005-12-31"))),
  mock_spec_survival("death", anchor = "entry",
    followup_min = 0, followup_max = 3650, event_prop = 0.3),
  mock_spec_survival("event", anchor = "entry",
    followup_min = 0, followup_max = 3650, event_prop = 0.5,
    censored_by = "death"),
  mock_spec_formula("death_days", formula = "as.numeric(death - entry)")
)

baseline <- generate_mock_data_native(spec, n = 1000, seed = 1)
dated <- generate_survival_dates(baseline, spec, seed = 1)
derived <- evaluate_mock_formulas(dated, spec, seed = 1)
names(baseline)
[1] "entry"
names(dated)
[1] "entry" "death" "event"
names(derived)
[1] "entry"      "death"      "event"      "death_days"

Each stage adds its own columns and each variable stays one column, so survival dates need no special handling in assembly, post-processing or diagnostics. Survival dates and formulas share one dependency field, depends_on, and one ordering routine: a variable is computed after the variables it depends on. Each stage draws from its own random-number stream, which the variables within that stage share. Adding a survival date therefore leaves baseline values unchanged, but two kinds of value can change: survival dates drawn after it, because all survival dates share one stream in dependency order; and the missing codes and garbage of later columns, because post-processing shares one stream across all columns. If you pin expected values in tests, regenerate them after adding a survival date:

entry <- mock_spec_date("entry", range = as.Date(c("2001-01-01", "2005-12-31")))
x <- mock_spec_continuous("x", range = c(0, 1),
  missing_codes = -99, missing_proportions = 0.3)
death_plain <- mock_spec_survival("death", anchor = "entry",
  followup_min = 0, followup_max = 10, event_prop = 0.5)
death_coded <- mock_spec_survival("death", anchor = "entry",
  followup_min = 0, followup_max = 10, event_prop = 0.5,
  missing_codes = "1900-01-01", missing_proportions = 0.3)

run_all <- function(s) {
  postprocess_mock_data(
    generate_survival_dates(generate_mock_data_native(s, n = 100, seed = 7), s, seed = 7),
    s, seed = 7
  )
}
without_survival <- mock_spec(entry, x)
generated_x_unchanged <- identical(
  generate_mock_data_native(without_survival, n = 100, seed = 7)$x,
  generate_mock_data_native(mock_spec(entry, death_coded, x), n = 100, seed = 7)$x
)
x_changed_plain <- sum(run_all(without_survival)$x != run_all(mock_spec(entry, death_plain, x))$x)
x_changed_coded <- sum(run_all(without_survival)$x != run_all(mock_spec(entry, death_coded, x))$x)

ltfu <- mock_spec_survival("ltfu", anchor = "entry",
  followup_min = 0, followup_max = 10, event_prop = 0.2)
clean_death <- function(s) {
  generate_survival_dates(generate_mock_data_native(s, n = 100, seed = 7), s, seed = 7)$death
}
count_changed <- function(a, b) sum(!mapply(identical, as.character(a), as.character(b)))
death_changed_before <- count_changed(clean_death(mock_spec(entry, death_plain)),
                                      clean_death(mock_spec(entry, ltfu, death_plain)))
death_changed_after <- count_changed(clean_death(mock_spec(entry, death_plain)),
                                     clean_death(mock_spec(entry, death_plain, ltfu)))

With the same seed, x is generated identically with or without the survival date. A new survival date drawn before death changes 70 of its 100 clean dates; one drawn after it changes 0. After post-processing, a survival date without missing codes changes 0 of the 100 returned x values, and one with missing codes changes 42, because it draws from the shared stream before x does.

How events are decided

Each survival date receives an event for exactly floor(n * event_prop) people, chosen at random, and is NA for everyone else. Follow-up times are whole days inside the window, spread by the chosen distribution (uniform, exponential or gompertz).

n_deaths <- sum(!is.na(dated$death))
death_days <- derived$death_days[!is.na(derived$death_days)]
range(death_days)
[1]   19 3643

Here 300 of 1,000 people have a death date, and every follow-up time is a whole number of days between 0 and 3,650.

A fixed count is the method of the legacy create_wide_survival_data(), kept in version 0.5 so that results match it. It suits mock data for testing pipelines. It cannot express a risk that differs from person to person, which a future hazard model would need.

The Gompertz parameters in the packaged minimal example (shape = 0.1, rate = 0.0001, time in days) show one limit of the ported method. The draws never exceed about three months, so every time is clamped to the start of the window (#54):

gompertz <- mock_spec(
  mock_spec_date("entry", range = as.Date(c("2001-01-01", "2005-12-31"))),
  mock_spec_survival("death", anchor = "entry",
    followup_min = 365, followup_max = 7300, event_prop = 1,
    distribution = "gompertz", shape = 0.1, rate = 1e-4)
)
gompertz_dates <- generate_survival_dates(
  generate_mock_data_native(gompertz, n = 500, seed = 1), gompertz, seed = 1
)
gompertz_days <- unique(as.numeric(gompertz_dates$death - gompertz_dates$entry))

All 500 deaths fall 365 days after entry, the lower end of a window that runs to 7,300 days.

Rules run on clean dates

After drawing a survival date, MockData applies two rules, in dependency order:

  1. A date earlier than its anchor becomes NA.
  2. If the variable has censored_by, it becomes NA wherever the censoring date is strictly earlier. A date censored by death is drawn after death, so it always compares against the final death date.

An event on the same day as the censoring date is kept:

ties <- mock_spec(
  mock_spec_date("entry", range = as.Date(c("2001-01-01", "2001-12-31"))),
  mock_spec_survival("death", anchor = "entry",
    followup_min = 20, followup_max = 20, event_prop = 1),
  mock_spec_survival("event", anchor = "entry",
    followup_min = 20, followup_max = 20, event_prop = 1,
    censored_by = "death")
)
tie_dates <- generate_survival_dates(
  generate_mock_data_native(ties, n = 100, seed = 1), ties, seed = 1
)
n_tied_events <- sum(!is.na(tie_dates$event))

All 100 events on the same day as death are kept. The rules run before post-processing, on the dates as generated. Missing codes and garbage come afterwards, so a garbage value added to test a cleaning pipeline, such as a date before entry, stays in the output instead of being removed by the rules.

Three layers of data

Because contamination comes last, generation produces three layers:

  • Clean truth: the generated dates and formula columns, before contamination.
  • Observed data: what postprocess_mock_data() and create_mock_data() return, after missing codes and garbage.
  • Analysis variables: status, follow-up time and similar quantities that an analysis recalculates from the observed data. MockData does not produce this layer.

A formula column belongs to the clean-truth layer, even though it is returned with the observed data. Here a death date is later shown as a missing code:

layers <- mock_spec(
  mock_spec_date("entry", range = as.Date(c("2001-01-01", "2001-12-31"))),
  mock_spec_survival("death", anchor = "entry",
    followup_min = 100, followup_max = 200, event_prop = 1,
    missing_codes = "1900-01-01", missing_proportions = 0.3),
  mock_spec_formula("death_days", formula = "as.numeric(death - entry)")
)
clean <- evaluate_mock_formulas(
  generate_survival_dates(
    generate_mock_data_native(layers, n = 200, seed = 2), layers, seed = 2
  ),
  layers, seed = 2
)
observed <- postprocess_mock_data(clean, layers, seed = 2)
shown_missing <- observed$death == as.Date("1900-01-01")
recomputed_days <- as.numeric(observed$death - observed$entry)

60 death dates are shown as the missing code 1900-01-01, but death_days is still between 100 and 200 on all of them, because it was computed from the clean dates. Recomputed from the returned dates, it would be about -37,095 days. For testing a cleaning pipeline that is what you want: the formula columns are the answer key. For analysis variables that must agree with the observed dates, compute them from the observed data after generation (vignette("survival-dates-v05") shows how).

Why chained censoring is rejected

A censored_by target may not have censored_by of its own. Consider three dates, A on day 30, B on day 20 and C on day 10, with A censored by B and B censored by C. Applied in order, the rules set B to NA, so A is no longer compared with an earlier date and would be reported as observed on day 30, although observation ended on day 10. MockData refuses the specification instead:

chained_message <- tryCatch(
  mock_spec(
    mock_spec_date("entry", range = as.Date(c("2001-01-01", "2001-12-31"))),
    mock_spec_survival("a", anchor = "entry", followup_min = 30,
      followup_max = 30, event_prop = 1, censored_by = "b"),
    mock_spec_survival("b", anchor = "entry", followup_min = 20,
      followup_max = 20, event_prop = 1, censored_by = "c"),
    mock_spec_survival("c", anchor = "entry", followup_min = 10,
      followup_max = 10, event_prop = 1)
  ),
  error = conditionMessage
)
cat(chained_message)
Survival variable 'a' has censored_by 'b', which is itself censored by 'c'. Chained censoring is not supported in this version; censor 'a' by the earliest date directly, or derive observed outcomes after generation.

A later version could keep every generated event time and derive observed outcomes from all censoring sources at once, which would allow several censoring sources per event.

What version 0.5 does not do

The stage order is fixed: survival dates are computed before formulas, so a formula can use a survival date, but a survival date cannot depend on a formula. An anchor must be a sampled date:

restriction_message <- tryCatch(
  mock_spec(
    mock_spec_date("entry", range = as.Date(c("2001-01-01", "2001-12-31"))),
    mock_spec_formula("entry_plus", formula = "entry"),
    mock_spec_survival("death", anchor = "entry_plus",
      followup_min = 0, followup_max = 10, event_prop = 1)
  ),
  error = conditionMessage
)
cat(restriction_message)
Survival variable 'death' has anchor 'entry_plus' of type 'formula'; an anchor must be a date variable.

This matters for exposure-dependent survival, where the hazard depends on an exposure such as smoking. The design leaves room for it, but some cases need more than new parameters:

Use case What would need adding Change to the design
Baseline smokers have a different hazard A survival model using each person’s exposure, with event occurrence following the model rather than a fixed count Fits the existing specification and stage
The hazard depends on derived pack-years Formulas computed before the survival dates that use them Scheduling across stages
Someone quits smoking during follow-up Exposure history, and a hazard that uses only the exposure known at each time A substantial generator extension
Illness changes smoking, which changes later outcomes Repeated updates to exposures and health states A longitudinal simulation mechanism

A causal reading of any of these would also need explicit assumptions about confounding and interventions, which the metadata does not yet express. Future survival models should keep the specification, the metadata adapters and the one-column-per-variable contract. They may need more parameters, scheduling across variable types, and separate rules for generating events and for observing them.