3  lnM effect sizes (SAFE)

Input Clean contrast data from Step 1.

Action Calculate SAFE lnM and its sampling variance.

Output Rdata/effect_sizes/proceed_lnm_safe.rds

Next Describe the effect-size dataset.

Code
source(here::here("Scripts", "00_packages.R"))
source(here::here("Scripts", "01_paths.R"))
source(here::here("Scripts", "04_lnm_safe.R"))
source(here::here("Scripts", "15_render_guard.R"))   # read-only: never recompute

# Set to TRUE to recompute (slow; ~hours for large datasets)
recompute_effect_sizes <- FALSE

3.1 Definition of lnM

This chapter computes \(\ln M\), the log ratio of the estimated between-group standard deviation to the pooled within-group standard deviation. The formula does not retain the direction of the difference between group means.

For a two-group comparison with observed summaries \((\bar{X}_1, s_1, n_1)\) and \((\bar{X}_2, s_2, n_2)\), the analytic definition begins with the one-way ANOVA mean squares:

\[ MS_B = \frac{n_1 n_2}{n_1 + n_2} (\bar{X}_1 - \bar{X}_2)^2, \]

\[ MS_W = \frac{(n_1 - 1)s_1^2 + (n_2 - 1)s_2^2} {n_1 + n_2 - 2}. \]

The within-group variance is \(s_W^2 = MS_W\). The between-component variance is:

\[ s_B^2 = \frac{MS_B - MS_W}{n_0}, \qquad n_0 = \frac{2n_1n_2}{n_1+n_2}, \]

where \(n_0\) is the harmonic-mean sample size. The effect size is then:

\[ \ln M = \ln\left(\frac{s_B}{s_W}\right), \]

where \(s_B\) is the square root of the between-component variance and \(s_W\) is the pooled within-group standard deviation. The sign of \(\ln M\) is therefore not the direction of trait change. Instead:

  • \(\ln M > 0\) means the estimated between-component standard deviation is larger than the within-group standard deviation.
  • \(\ln M = 0\) means the two estimated components have a ratio of one.
  • \(\ln M < 0\) means the estimated between-component standard deviation is smaller than the within-group standard deviation.

For large, balanced independent groups, \(\ln M = 0\) corresponds approximately to \(d = \sqrt{2}\), or about 1.4, on the standardized mean-difference scale.

This calculation differs from taking an absolute value of a signed effect size such as \(|d|\) or \(|\ln RR|\). The SAFE code below calculates \(\ln M\) directly.

3.2 SAFE calculation

The analytic calculation has two numerical limits:

  • Sampling variance can become very large when \(MS_B\) is close to \(MS_W\).
  • The plug-in estimate is undefined when \(MS_B \le MS_W\), because \(s_B^2\) is then zero or negative.

The SAFE estimator repeatedly draws values of the group means and standard deviations from their finite-sample distributions. It transforms each draw with \(MS_B^\ast > MS_W^\ast\) to \(\ln M^\ast\), then saves the bias-corrected estimate and sampling variance.

The book uses the cached SAFE results by default. Recomputing all contrasts is opt-in. The Monte Carlo step is slow.

3.3 Load cleaned data

Code
dat_lnm_input <- readRDS(
  here::here("Rdata", "data_clean", "proceed_clean_filtered.rds")
)

cat("Contrasts available for lnM:", nrow(dat_lnm_input), "\n")
Contrasts available for lnM: 7250 

3.4 Compute or read SAFE lnM estimates

Code
dat_es <- compute_or_read_lnm_safe(
  dat_lnm_input,
  out_path   = here::here("Rdata", "effect_sizes", "proceed_lnm_safe.rds"),
  diag_path  = here::here("Rdata", "tables", "safe_lnm_diagnostics.csv"),
  recompute  = recompute_effect_sizes
)

cat("Effect sizes retained:", nrow(dat_es), "\n")
Effect sizes retained: 7237 

The resulting dataset carries the SAFE point estimate (yi_lnM_safe) and its sampling variance (vi_lnM_safe). These two columns become the response and known sampling-error input for the location-scale meta-regression chapters.

3.5 Effect-size summary

Code
dat_es |>
  dplyr::summarise(
    k           = dplyr::n(),
    min_yi      = round(min(yi_lnM_safe), 3),
    median_yi   = round(median(yi_lnM_safe), 3),
    max_yi      = round(max(yi_lnM_safe), 3),
    min_vi      = round(min(vi_lnM_safe), 6),
    median_vi   = round(median(vi_lnM_safe), 6),
    max_vi      = round(max(vi_lnM_safe), 6),
    n_status_ok = sum(status == "ok", na.rm = TRUE),
    n_failed    = sum(status != "ok", na.rm = TRUE)
  )

This table records the median and range of yi_lnM_safe and vi_lnM_safe.

3.6 SAFE diagnostics

SAFE summarizes valid simulation draws. The retained percentage records the share of draws for which the between-group component was positive.

Code
safe_diag <- readr::read_csv(
  here::here("Rdata", "tables", "safe_lnm_diagnostics.csv"),
  show_col_types = FALSE
)

safe_diag |>
  dplyr::summarise(
    n_ok           = sum(status == "ok"),
    n_not_ok       = sum(status != "ok"),
    min_pct_kept   = round(min(draws_kept / draws_total * 100), 1),
    median_pct_kept = round(median(draws_kept / draws_total * 100), 1),
    max_attempts   = max(attempts)
  )

3.7 Identifier management

The dataset preserves two IDs:

  • es_id_db: the original PROCEED effect-size identifier, kept for tracing estimates back to the database.
  • es_id_model: a sequential factor created after filtering to finite, positive \(v_i\), used as the row/column index of the sampling-variance matrix \(V\).
Code
dat_es |>
  dplyr::select(es_id_db, es_id_model) |>
  head(6)

3.8 Sampling-variance matrix

The diagonal covariance matrix \(V\) carries the known SAFE sampling variances into brms via (1 | gr(es_id_model, cov = V)). The corresponding standard deviation prior is fixed with constant(1), so this term contributes known measurement error rather than estimating a new variance component.

Code
cat("Effect-size rows (= dim of V):", nrow(dat_es), "\n")
Effect-size rows (= dim of V): 7237