Changelog
Source:NEWS.md
MockData 0.5.0 (2026-09-27)
Breaking changes
-
Seeded output changed once for the RNG-stream mechanism in this release. Within the
create_mock_data()pipeline (generate_mock_data_native(),generate_mock_data_simstudy(),postprocess_mock_data()), MockData now derives all randomness from independent L’Ecuyer-CMRG sub-streams (one per generation stage) seeded from the single publicseed, replacing the previous Mersenne-Twisterseed/seed + 1scheme. The standalonecreate_*helpers (create_cat_var(),create_con_var(),create_date_var(),create_survival_dates(),create_wide_survival_data(),sample_with_proportions(),make_garbage(),apply_garbage()) called directly are unaffected and still seed a single Mersenne-Twister stream per call. For a given seed and package version pipeline output is reproducible and independent of the session’s ambientRNGkind(); it is not comparable across the v0.4 → v0.5 boundary. For seeded calls, generation no longer alters the caller’s RNG state (previously the legacyvalidate = FALSEpath reset it). (#38)Migrating: if your tests pin values from seeded v0.4 runs of
create_mock_data()or themock_specpipeline, regenerate those expected values once under v0.5 and pin them again. The same seed keeps giving identical output within this version. Standalonecreate_*calls need no change. Migration: add
anchor(andcensored_bywhere a competing risk applies) to each survival date’s row. A date withfollowup_min,followup_maxorevent_propbut noanchornow fails with a message naming the fix; previously it was silently dropped or generated as a plain calendar date.Missing codes and garbage are applied to survival dates after the survival rules. Garbage that places a date before entry is therefore kept, whereas the legacy engine set such dates to
NA.The packaged minimal example now generates entirely through the v0.4 pipeline, because its Gompertz survival dates no longer force the legacy generator. Its seeded output differs from v0.4.
If the legacy generator is selected (
validate = FALSE,variable_details = NULL, detail-level-onlydatabaseStart, or another unsupported variable) on metadata that setsanchor,create_mock_data()stops and names the survival variables rather than dropping them.The same applies to formula variables. If the legacy generator is selected for metadata that sets
mockFormulaon an enabled variable,create_mock_data()stops and names the formula variables. Previously the legacy generator ignored the formula and returned unrelated random values, or dropped aDerivedVar::variable, without saying so.On the legacy path (
validate = FALSEand the other legacy triggers),create_mock_data()now generates only variables whose variable-leveldatabaseStartincludes the requested database, as the v0.4 pipeline does. Previously it chose variables from the detail rows alone, so a variable listed only for another database was generated as random values, and a formula or survival variable belonging to another database stopped the run.
New features
- Formula-derived variables (#39, Phase A). Variables can now be generated from algebraic expressions over other generated variables via a new
mockFormulacolumn invariable_details(e.g.weight / (height^2)), or the newmock_formula()direct API. Formulas are evaluated in dependency order byevaluate_mock_formulas(), in a restricted environment exposing only the generated columns and a fixed allow-list of base functions.DerivedVar::/Func::semantics are unchanged: derived variables without amockFormularemain excluded from generation, andFunc::dispatch is not yet supported. -
create_mock_data()now generates survival dates from metadata. A date whosevariables.csvrow setsanchor(its entry-date variable) is a survival date:floor(n * event_prop)rows receive a date betweenfollowup_minandfollowup_maxdays after the anchor, drawn from a uniform, exponential or Gompertz distribution, and the rest areNA. An optionalcensored_bycolumn names a competing survival date, such as death, that sets this date toNAwhere it comes first. The statistics and rules are those ofcreate_wide_survival_data(), ported exactly. New functions:mock_spec_survival()andgenerate_survival_dates(). This resolves the known issue, listed since v0.2.0, that survival data had to be generated separately. - Derived columns, from
mockFormulaor survival dates, describe the clean generated values; missing codes and garbage are then applied to each column independently.is.nais added to themockFormulaallow-list so status and follow-up time can be derived. Derive both from the same observation window:as.integer(!is.na(death_date))only says a death date exists, not that the death was observed before censoring. -
create_mock_data()now acceptsn = 0, returning a zero-row data frame with the full generated schema (useful for schema tests), and rejects fractional, negative,NA, and non-finite values ofnwith a clear message. The validation now matchesgenerate_mock_data_native(). - The native backend now supports
distribution = "exponential"(parity with the legacy generator), removing a forced legacy-fallback for exponential metadata. When the optional simstudy backend is selected, exponential variables are routed to the native generator (like all non-uniform continuous distributions); the simstudy package itself is not involved in their generation. Native exponential values are truncated to the declaredrangeby inverse-CDF sampling, rather than clipped at the range maximum (with a point mass at the boundary) as the legacyrexp()path does.
Validation and error messages
Calling
create_mock_data()withoutdatabaseStartnow fails upfront with a message naming the argument, instead of a raw missing-argument error.With
validate = TRUE(the default), invalid distribution parameters in metadata — e.g.distribution = "exponential"without a positiverate— now stop generation with a message naming the variable and how to fix it, instead of warning and substituting a uniform draw. The legacy warn-and-substitute behaviour remains available viavalidate = FALSE.A
censored_bytarget may not itself havecensored_by. Chained censoring would report some events as observed after observation ended, so it is rejected in this version with a message naming the fix.
Bug fixes
- The simstudy backend now returns a typed zero-row data frame for
n = 0, identical to the native backend’s output, instead of failing insidesimstudy::genData()(whose1:nid table has two rows whenn = 0). (#50) - The v0.4 pipeline now expands range-notation missing codes such as
[997,999](the usual CCHS and CHMS pattern for don’t know, refusal and not stated) into their individual codes, splitting the row’s proportion equally. Previously a continuous variable with such a code failed in post-processing, and a categorical variable silently wrote the literal string"[997,999]"into the data. A bracketed missing code that is not an integer range now fails with a message naming the variable. (#58) - The packaged minimal example corrects two metadata errors: BMI’s low garbage range had a stray parenthesis (
[-10;15])), and height’s low-garbage proportion was 1.00 (every row) instead of 0.01; height’s high-garbage range had an infinite upper bound,(2.1;inf], which producedNaNvalues, and is now(2.1;2.5](#60). With these, #58 and the survival dates (#40), every variable in the example generates through the v0.4 pipeline. Seeded output for the example changes.
Deprecations
-
create_wide_survival_data()is deprecated and warns once per session.
MockData 0.4.0 (2026-06-10)
Breaking changes
-
create_cat_var(),create_con_var(), andcreate_date_var()now stop with the errorVariable '<name>' not found in variables metadatawhen the requested variable is absent, instead of warning and returningNULL. This affects direct generator calls andcreate_wide_survival_data()(a misspelled date-variable name now errors instead of being skipped with a warning).create_mock_data()itself derives variable names from thevariablesmetadata, so it cannot trigger this error; itsvalidate = FALSEflag continues to convert any generator error to warn-and-skip on the legacy path. Duplicatevariablesrows for the same variable now produce a warning in the legacycreate_*path before the first row is used (the v0.4mock_specpath already errors on duplicate names). -
create_mock_data()error messages for missing metadata files changed fromConfiguration file does not exist:/Details file does not exist:tovariables file does not exist:/variable_details file does not exist:.
New features
- Started the v0.4 production refactor around a normalized
mock_specarchitecture. - Added
mock_spec(),mock_spec_continuous(),mock_spec_categorical(),mock_spec_date(),is_mock_spec(), andvalidate_mock_spec(). - Added direct specification helpers
mock_continuous(),mock_categorical(), andmock_date()for simple use without recodeflow-style metadata tables. - Added
mock_spec_from_recodeflow()to adapt recodeflow-stylevariablesandvariable_detailsmetadata into validatedmock_specobjects while preserving role/database filtering, categorical proportions,recEndmissing-code semantics, valid ranges, garbage rules, date ranges, and survival/date fields. - Added
generate_mock_data_native()to generate baseline valid mock data frommock_specobjects with the native R backend. - Added
postprocess_mock_data()to applymock_specmissing-code and garbage-value rules after baseline generation, with diagnostics that distinguish assigned missing/garbage rows from naturally drawn values. - Post-processing diagnostics now protect naturally drawn missing-code collisions from later garbage assignment, apply garbage rules in canonical
low->high-> other order, and stop on repeated post-processing. This prevents silent diagnostic drift when a naturally drawn missing-code value would otherwise be overwritten by garbage assignment. - Added
generate_mock_data_simstudy()as a soft-gated optional backend for baseline categorical and uniform continuous generation whensimstudyis installed, with native generation retained for MockData-specific semantics. - The optional
simstudybackend is kept inSuggests, requiressimstudy >= 0.8.1, and validates categorical labels before converting generated values back into MockData’smock_speclevels. - The optional
simstudybackend now rejects variables namedid, which conflicts withsimstudy’s generated row identifier, and normalizes categorical output through an explicit label-or-index validation path. -
create_mock_data()now attempts the v0.4mock_specpipeline in strict mode for supported recodeflow metadata, while retaining the legacycreate_*dispatch path for unsupported v0.4 backend features and lenient generation. The v0.4 path attachesmockdata_diagnosticsand usesseedfor baseline generation plusseed + 1for post-processing, so exact seeded output may differ from v0.3.x even when the public seed is unchanged. Verbose mode now reports whether the v0.4 or legacy path was chosen. - Added forward-compatible specification fields:
spec_version,provenance, andmodel_hint. - Existing v0.3 generator APIs remain available while v0.4 internals are built.
MockData 0.3.0
Breaking changes
New function API - All generator functions now accept full metadata data frames instead of pre-filtered subsets:
# Before (v0.2.x)
var_row <- variables[variables$variable == "age", ]
details_subset <- variable_details[variable_details$variable == "age", ]
result <- create_con_var(var_row, details_subset, n = 1000)
# After (v0.3.0)
result <- create_con_var(
var = "age",
databaseStart = "minimal-example",
variables = variables,
variable_details = variable_details,
n = 1000
)Affected functions: create_cat_var(), create_con_var(), create_date_var(), create_wide_survival_data(), create_mock_data()
Deprecated: prop_garbage parameter in create_wide_survival_data(). Use garbage parameters in metadata instead:
# Old way (no longer supported)
surv <- create_wide_survival_data(..., prop_garbage = 0.03)
# New way
vars_with_garbage <- add_garbage(variables, "death_date",
garbage_high_prop = 0.03, garbage_high_range = "[2025-01-01, 2099-12-31]")
surv <- create_wide_survival_data(..., variables = vars_with_garbage)New features
Unified garbage generation across all variable types (categorical, continuous, date, survival):
-
garbage_low_prop+garbage_low_rangefor values below valid range -
garbage_high_prop+garbage_high_rangefor values above valid range - New helper function
add_garbage()for easy garbage specification - Categorical garbage now supported (treats codes as ordinal to generate out-of-range values)
# Add garbage to any variable type
vars_with_garbage <- add_garbage(variables, "smoking",
garbage_low_prop = 0.02, garbage_low_range = "[-2, 0]")
# Pipe-friendly
vars_with_garbage <- variables %>%
add_garbage("age", garbage_high_prop = 0.03, garbage_high_range = "[150, 200]") %>%
add_garbage("smoking", garbage_low_prop = 0.02, garbage_low_range = "[-2, 0]")Derived variable identification:
-
identify_derived_vars()- Identifies derived variables usingDerivedVar::andFunc::patterns -
get_raw_var_dependencies()- Extracts raw variable dependencies - Compatible with recodeflow patterns
Bug fixes
- Fixed categorical garbage factor level bug - garbage values were being converted to NA during factor creation
- Fixed
recEndcolumn requirement - now optional for simple configurations - Fixed derived variable generation in
create_mock_data()- derived variables now correctly excluded - Fixed
create_mock_data()error handling: strict mode is now the default, so unsupportedrTypevalues and generator errors stop generation instead of silently dropping columns. Usevalidate = FALSEto opt into warning-and-skip behavior. - Fixed role matching so
disabledno longer matchesenabled. - Fixed missing-code classification to use
recEndmetadata instead of numeric-code heuristics, so valid codes such as 7, 17, and 27 are not misclassified as missing. - Fixed
variable_details = NULLfallback mode for categorical, continuous, and date variables. - Fixed
rTypenormalization sodate/Datehandling is consistent across validation, defaults, and generation. - Added migration warnings for legacy
corrupt_*garbage fields and recStart values, rewriting them to canonicalgarbage_*names. - Added validation that
garbage_low_prop + garbage_high_propdoes not exceed 1, preventing silent truncation of the second garbage pass.
Documentation
Restructured using Divio framework:
- Removed 6 vignettes (cchs-example, chms-example, demport-example, dates, schema-change-dates, tutorial-config-files)
- Added 2 new vignettes (tutorial-categorical-continuous, tutorial-survival-data)
- Massively expanded reference-config (2,028 lines of comprehensive metadata schema documentation)
- All vignettes updated to v0.3.0 API
- All examples now use
inst/extdata/minimal-example/only
Final structure (9 vignettes):
- Tutorials (6): getting-started, tutorial-categorical-continuous, tutorial-dates, tutorial-survival-data, tutorial-missing-data, tutorial-garbage-data
- How-to guides (1): for-recodeflow-users
- Explanation (1): advanced-topics
- Reference (1): reference-config
Metadata simplification:
- Removed
inst/extdata/cchs/,inst/extdata/chms/,inst/extdata/demport/ - Only
inst/extdata/minimal-example/remains as canonical reference
Migration guide
Update function calls:
- Pass variable name as string (not pre-filtered row)
- Pass full metadata data frames (not subsets)
- Add
databaseStartparameter - Remove manual filtering
Update garbage specification:
- Remove
prop_garbagefromcreate_wide_survival_data()calls - Add garbage to metadata using
add_garbage()helper
MockData 0.2.0
Major changes
New configuration format (v0.2)
-
Breaking change: New configuration schema with
uid/uid_detailsystem - Replaces v0.1
cat/catLabelcolumns with unified metadata structure - Adds
rTypefield for explicit R type coercion (factor, integer, double, Date) - Adds
proportionfield for direct distribution control - Adds date-specific fields:
date_start,date_end,distribution
Backward compatibility: v0.1 format still supported via dual interface. Both formats work side-by-side.
Date variable generation
- New
create_date_var()function for date variables - Multiple distribution options: uniform, gompertz, exponential
- Support for survival analysis patterns
- SAS date format parsing
- Three source formats: analysis (R Date), csv (ISO strings), sas (numeric)
Survival analysis support
- New
create_wide_survival_data()function for cohort studies - Generates paired entry and event dates with guaranteed temporal ordering
- Supports censoring and multiple event distributions
-
Note: Must be called manually (not compatible with
create_mock_data()batch generation)
Data quality testing (garbage data)
- New
prop_invalidparameter across all generators - Generates intentionally invalid data for testing validation pipelines
- Supports garbage types:
corrupt_future,corrupt_past,corrupt_range - Critical for testing data cleaning workflows
Batch generation
- New
create_mock_data()function for batch generation from CSV configuration - New
read_mock_data_config()andread_mock_data_config_details()readers - Processes multiple variables in single call
- Fallback mode when details not provided
New functions
-
create_date_var()- Date variable generation -
create_wide_survival_data()- Paired survival dates with temporal ordering -
create_mock_data()- Batch generation orchestrator -
read_mock_data_config()- Configuration file reader -
read_mock_data_config_details()- Details file reader -
determine_proportions()- Unified proportion determination -
import_from_recodeflow()- Helper to adapt recodeflow metadata
Function updates
-
create_cat_var(): Add rType support, proportion parameter, uid-based filtering -
create_con_var(): Add rType support, proportion parameter for missing codes - Consolidate helpers in
mockdata_helpers.R,config_helpers.R,scalar_helpers.R
Documentation
New vignettes
-
getting-started.qmd- Comprehensive introduction -
tutorial-dates.qmd- Date configuration patterns -
tutorial-config-files.qmd- Batch generation workflow -
reference-config.qmd- Complete v0.2 schema documentation -
advanced-topics.qmd- Technical implementation details
Package infrastructure
- Added
_pkgdown.ymlfor documentation website - Updated NAMESPACE with new imports (stats::rexp, utils::read.csv, etc.)
- Updated DESCRIPTION with new dependencies
Breaking changes
Configuration format changes:
- Variable details now require
uidanduid_detailcolumns -
rTypefield required for proper type coercion - New date fields:
date_start,date_end,distribution
Migration path:
- v0.1 format still works (backward compatibility maintained)
- Dual interface auto-detects format based on parameters
- v0.2 recommended for new projects
File changes:
- Renamed
R/mockdata-helpers.R→R/mockdata_helpers.R - ICES metadata removed (maintained in recodeflow package)
Bug fixes
- Fixed ‘else’ handling in
recEndrules (issue #5) - Fixed create_wide_survival_data() compatibility with create_mock_data()
- Fixed Roxygen documentation link syntax errors
Known issues
- Survival variable type must be generated manually with
create_wide_survival_data()(resolved in 0.5.0, #40) - Cannot be used in
create_mock_data()batch generation (requires paired variables) (resolved in 0.5.0, #40)