Action Calculate SAFE lnM and its sampling variance.
OutputRdata/effect_sizes/proceed_lnm_safe.rds
Next Describe the effect-size dataset.
Code
source(here::here("Scripts", "00_packages.R"))source(here::here("Scripts", "01_paths.R"))source(here::here("Scripts", "04_lnm_safe.R"))source(here::here("Scripts", "15_render_guard.R")) # read-only: never recompute# Set to TRUE to recompute (slow; ~hours for large datasets)recompute_effect_sizes <-FALSE
3.1 Definition of lnM
This chapter computes \(\ln M\), the log ratio of the estimated between-group standard deviation to the pooled within-group standard deviation. The formula does not retain the direction of the difference between group means.
For a two-group comparison with observed summaries \((\bar{X}_1, s_1, n_1)\) and \((\bar{X}_2, s_2, n_2)\), the analytic definition begins with the one-way ANOVA mean squares:
where \(n_0\) is the harmonic-mean sample size. The effect size is then:
\[
\ln M = \ln\left(\frac{s_B}{s_W}\right),
\]
where \(s_B\) is the square root of the between-component variance and \(s_W\) is the pooled within-group standard deviation. The sign of \(\ln M\) is therefore not the direction of trait change. Instead:
\(\ln M > 0\) means the estimated between-component standard deviation is larger than the within-group standard deviation.
\(\ln M = 0\) means the two estimated components have a ratio of one.
\(\ln M < 0\) means the estimated between-component standard deviation is smaller than the within-group standard deviation.
For large, balanced independent groups, \(\ln M = 0\) corresponds approximately to \(d = \sqrt{2}\), or about 1.4, on the standardized mean-difference scale.
This calculation differs from taking an absolute value of a signed effect size such as \(|d|\) or \(|\ln RR|\). The SAFE code below calculates \(\ln M\) directly.
3.2 SAFE calculation
The analytic calculation has two numerical limits:
Sampling variance can become very large when \(MS_B\) is close to \(MS_W\).
The plug-in estimate is undefined when \(MS_B \le MS_W\), because \(s_B^2\) is then zero or negative.
The SAFE estimator repeatedly draws values of the group means and standard deviations from their finite-sample distributions. It transforms each draw with \(MS_B^\ast > MS_W^\ast\) to \(\ln M^\ast\), then saves the bias-corrected estimate and sampling variance.
The book uses the cached SAFE results by default. Recomputing all contrasts is opt-in. The Monte Carlo step is slow.
3.3 Load cleaned data
Code
dat_lnm_input <-readRDS( here::here("Rdata", "data_clean", "proceed_clean_filtered.rds"))cat("Contrasts available for lnM:", nrow(dat_lnm_input), "\n")
The resulting dataset carries the SAFE point estimate (yi_lnM_safe) and its sampling variance (vi_lnM_safe). These two columns become the response and known sampling-error input for the location-scale meta-regression chapters.
es_id_db: the original PROCEED effect-size identifier, kept for tracing estimates back to the database.
es_id_model: a sequential factor created after filtering to finite, positive \(v_i\), used as the row/column index of the sampling-variance matrix \(V\).
The diagonal covariance matrix \(V\) carries the known SAFE sampling variances into brms via (1 | gr(es_id_model, cov = V)). The corresponding standard deviation prior is fixed with constant(1), so this term contributes known measurement error rather than estimating a new variance component.
Code
cat("Effect-size rows (= dim of V):", nrow(dat_es), "\n")
---title: "lnM effect sizes (SAFE)"---::: {.workflow}::: {}**Input**Clean contrast data from Step 1.:::::: {}**Action**Calculate SAFE lnM and its sampling variance.:::::: {}**Output**`Rdata/effect_sizes/proceed_lnm_safe.rds`:::::: {}**Next**Describe the effect-size dataset.::::::```{r setup}source(here::here("Scripts", "00_packages.R"))source(here::here("Scripts", "01_paths.R"))source(here::here("Scripts", "04_lnm_safe.R"))source(here::here("Scripts", "15_render_guard.R")) # read-only: never recompute# Set to TRUE to recompute (slow; ~hours for large datasets)recompute_effect_sizes <-FALSE```## Definition of lnM {#sec-lnm-measures}This chapter computes $\ln M$, the log ratio of the estimated between-groupstandard deviation to the pooled within-group standard deviation. The formuladoes not retain the direction of the difference between group means.For a two-group comparison with observed summaries $(\bar{X}_1, s_1, n_1)$ and $(\bar{X}_2, s_2, n_2)$, the analytic definition begins with the one-way ANOVA mean squares:$$MS_B =\frac{n_1 n_2}{n_1 + n_2}(\bar{X}_1 - \bar{X}_2)^2,$$$$MS_W =\frac{(n_1 - 1)s_1^2 + (n_2 - 1)s_2^2}{n_1 + n_2 - 2}.$$The within-group variance is $s_W^2 = MS_W$. The between-component variance is:$$s_B^2 =\frac{MS_B - MS_W}{n_0},\qquadn_0 = \frac{2n_1n_2}{n_1+n_2},$$where $n_0$ is the harmonic-mean sample size. The effect size is then:$$\ln M = \ln\left(\frac{s_B}{s_W}\right),$$where $s_B$ is the square root of the between-component variance and $s_W$ is the pooled within-group standard deviation. The sign of $\ln M$ is therefore not the direction of trait change. Instead:- $\ln M > 0$ means the estimated between-component standard deviation is larger than the within-group standard deviation.- $\ln M = 0$ means the two estimated components have a ratio of one.- $\ln M < 0$ means the estimated between-component standard deviation is smaller than the within-group standard deviation.For large, balanced independent groups, $\ln M = 0$ corresponds approximatelyto $d = \sqrt{2}$, or about 1.4, on the standardized mean-difference scale.This calculation differs from taking an absolute value of a signed effect sizesuch as $|d|$ or $|\ln RR|$. The SAFE code below calculates $\ln M$ directly.## SAFE calculationThe analytic calculation has two numerical limits:- Sampling variance can become very large when $MS_B$ is close to $MS_W$.- The plug-in estimate is undefined when $MS_B \le MS_W$, because $s_B^2$ is then zero or negative.The SAFE estimator repeatedly draws values of the group means and standarddeviations from their finite-sample distributions. It transforms each draw with$MS_B^\ast > MS_W^\ast$ to $\ln M^\ast$, then saves the bias-corrected estimateand sampling variance.The book uses the cached SAFE results by default. Recomputing all contrasts is opt-in. The Monte Carlo step is slow.## Load cleaned data```{r load-clean}dat_lnm_input <-readRDS( here::here("Rdata", "data_clean", "proceed_clean_filtered.rds"))cat("Contrasts available for lnM:", nrow(dat_lnm_input), "\n")```## Compute or read SAFE lnM estimates```{r compute-lnm}dat_es <-compute_or_read_lnm_safe( dat_lnm_input,out_path = here::here("Rdata", "effect_sizes", "proceed_lnm_safe.rds"),diag_path = here::here("Rdata", "tables", "safe_lnm_diagnostics.csv"),recompute = recompute_effect_sizes)cat("Effect sizes retained:", nrow(dat_es), "\n")```The resulting dataset carries the SAFE point estimate (`yi_lnM_safe`) and its sampling variance (`vi_lnM_safe`). These two columns become the response and known sampling-error input for the location-scale meta-regression chapters.## Effect-size summary```{r es-summary}dat_es |> dplyr::summarise(k = dplyr::n(),min_yi =round(min(yi_lnM_safe), 3),median_yi =round(median(yi_lnM_safe), 3),max_yi =round(max(yi_lnM_safe), 3),min_vi =round(min(vi_lnM_safe), 6),median_vi =round(median(vi_lnM_safe), 6),max_vi =round(max(vi_lnM_safe), 6),n_status_ok =sum(status =="ok", na.rm =TRUE),n_failed =sum(status !="ok", na.rm =TRUE) )```This table records the median and range of `yi_lnM_safe` and`vi_lnM_safe`.## SAFE diagnosticsSAFE summarizes valid simulation draws. The retained percentage records theshare of draws for which the between-group component was positive.```{r safe-diag}safe_diag <- readr::read_csv( here::here("Rdata", "tables", "safe_lnm_diagnostics.csv"),show_col_types =FALSE)safe_diag |> dplyr::summarise(n_ok =sum(status =="ok"),n_not_ok =sum(status !="ok"),min_pct_kept =round(min(draws_kept / draws_total *100), 1),median_pct_kept =round(median(draws_kept / draws_total *100), 1),max_attempts =max(attempts) )```## Identifier managementThe dataset preserves two IDs:- `es_id_db`: the original PROCEED effect-size identifier, kept for tracing estimates back to the database.- `es_id_model`: a sequential factor created after filtering to finite, positive $v_i$, used as the row/column index of the sampling-variance matrix $V$.```{r id-check}dat_es |> dplyr::select(es_id_db, es_id_model) |>head(6)```## Sampling-variance matrixThe diagonal covariance matrix $V$ carries the known SAFE sampling variances into `brms` via `(1 | gr(es_id_model, cov = V))`. The corresponding standard deviation prior is fixed with `constant(1)`, so this term contributes known measurement error rather than estimating a new variance component.```{r vcv-dims}cat("Effect-size rows (= dim of V):", nrow(dat_es), "\n")```