Learning objectives

By the end of this session, you will be able to:

  1. Explain what an effect size is and why each one comes with a sampling variance.
  2. Choose the right metric for your data: mean difference, Hedges’ g, log response ratio, log odds ratio, Fisher’s z.
  3. Compute them by hand and with metafor::escalc().
  4. Recognise the traps: small samples, near-zero means, missing SDs, unsafe conversions.

Studies do not speak the same language

The 888 intercropping comparisons of our running example report biodiversity as:

  • abundance or species richness,
  • in individuals per trap, per plant, per m², per hour…
  • with SD, SE, 95% CI or median + IQR as the measure of spread.

An effect size puts all studies on one common scale: the direction and magnitude of the difference between intervention and comparator — plus its uncertainty.

Source: Jones et al. (2021); the authors converted everything to SD before sharing the data.

Why not just use p-values?

A p-value mixes effect size and sample size: a tiny effect can be “highly significant”, a large one “not significant”.

Meta-analysis combines effect sizes and their uncertainty, not p-values (Nakagawa and Cuthill 2007).

Every study sees the truth through noise

θ̂k=θk+εkobserved=true+sampling error,Var(εk)=vk\hat\theta_k = \theta_k + \varepsilon_k \qquad \text{observed} = \text{true} + \text{sampling error}, \quad \text{Var}(\varepsilon_k) = v_k

Each effect size comes with its sampling variance vkv_k. In the meta-analysis this afternoon, each study is weighted by the inverse of its variance, wk=1/vkw_k = 1/v_k: precise studies count more. (In a random-effects model, the weight also includes the real variation between studies, τ2\tau^2: wk=1/(vk+τ2)w_k = 1/(v_k + \tau^2).)

A decision tree

Continuous outcome, two groups
means, SD, n
  • Same unit everywhere → Mean difference (MD)
  • Different scales, additive effect → standardised mean difference (Hedges' g)
  • Positive values, proportional effect → log response ratio (lnRR)
Binary outcome, two groups
counts in a 2×2 table
  • Presence/absence, survival, infestation → log odds ratio (or log risk ratio)
Relationship between two continuous variables
correlation r, n
  • e.g. yield vs landscape complexity → Fisher's z

Proportional (→ lnRR): “−30% aphids”, same meaning with 10 or 1,000 aphids. Abundance, biomass, yield: a true zero, effects multiply.
Additive (→ g): “+5 units”. Scores, indices, variables with an arbitrary zero.

In ecology and agronomy, lnRR and Hedges’ g (a standardised mean difference, SMD) dominate. Our running example → lnRR. Fix the metric in the protocol, before seeing the results.

Show of hands: which metric for your own data?

Our worked example

Aphids on fennel, intercropped with cotton vs fennel monoculture (Fernandes et al. 2013), one comparison of our dataset:

Intercropping (T) Monoculture (C)
Mean 8.37 23.98
SD 2.85 5.46
n 4 4

Convention for the whole week: T (intervention) in the numerator / first, C (comparator) second. A negative effect = fewer pests with intercropping.

The log response ratio (lnRR)

lnRR=ln(X‾TX‾C)\ln RR = \ln\left(\frac{\bar X_T}{\bar X_C}\right)

vlnRR=SDT2nTX‾T2+SDC2nCX‾C2v_{\ln RR} = \frac{SD_T^2}{n_T \bar X_T^2} + \frac{SD_C^2}{n_C \bar X_C^2}

(Hedges et al. 1999)

Back to percent for reporting:

change=100×(elnRR−1)%\text{change} = 100 \times \left(e^{\ln RR} - 1\right)\%

Why the log? The ratio is asymmetric (halving = 0.5, doubling = 2); its log is symmetric (−0.69, +0.69) and closer to normal.

Why ecologists like it Unit-free and proportional: “intercropping reduces aphids by 65%” speaks to agronomists — and its value does not depend on the SD (only its variance does).

Exercise 1 · By hand (pairs, 8 min)

Calculator ready

With the aphid example (T: 8.37 ± 2.85, n = 4; C: 23.98 ± 5.46, n = 4), compute lnRR, its variance, and the % change.

lnRR = ln(8.37 / 23.98) = -1.053
v = 2.85² / (4 × 8.37²) + 5.46² / (4 × 23.98²) = 0.0419 (SE = 0.205)
Change = 100 × (e-1.053 − 1) = -65%

Now let R do it (live coding, 5 min)

Predict first What value of yi should R return for our aphid example?

library(metafor)
# ic <- read.csv("intercropping_biodiversity.csv")      # running example

dat <- escalc(measure = "ROM",                           # lnRR ("ratio of means")
              m1i = b_mean_t, sd1i = b_sd_t, n1i = b_n_t,  # T = intercropping
              m2i = b_mean_c, sd2i = b_sd_c, n2i = b_n_c,  # C = monoculture
              data = ic)
# other metrics: measure = "SMD" (Hedges' g), "OR", "ZCOR"

dat[dat$es_id == 2011, c("es_id", "yi", "vi")]          # our worked example

    es_id      yi     vi 
137  2011 -1.0526 0.0419 
c(NA_yi = sum(is.na(dat$yi)), vi_zero = sum(dat$vi == 0, na.rm = TRUE))  # what to deal with
  NA_yi vi_zero 
    156      46 
  • yi = effect size, vi = sampling variance: the input of every meta-analysis (rma(), this afternoon).
  • Zero mean → NA; zero SD → vi = 0 (“infinite” precision!): count both before rma().

Hedges’ g, step by step

Idea: express the difference in standard-deviation units.

d=X‾T−X‾CSDpooledd = \frac{\bar X_T - \bar X_C}{SD_{pooled}}

Useful when studies use different scales and the effect is additive.

Small samples inflate d → multiply by a correction J < 1: that is Hedges’ g (Hedges 1981).

With our aphids:

1. Difference: 8.37 − 23.98 = -15.61 aphids

2. Pooled SD (equal n: average the variances): √((2.85² + 5.46²) / 2) = 4.36

3. d = -15.61 / 4.36 = -3.58 standard deviations

4. With 4 + 4 plants, J = 0.87 → g = -3.12

escalc(measure = "SMD") does all of this — and the variance. You never compute g by hand in practice.

Trap 1 · Means close to zero

Your call

156 comparisons in our data have a zero mean. What do you do with them for a lnRR?

Exact zeros: no lnRR (and most also have SD = 0, so no g). Report them; test the conclusion with and without.

Near zero: lnRR is biased. Geary’s rule: each mean ≥ 3 SE from zero → 28% of our usable comparisons fail it.

What to do

  • Near zero or small n: bias-corrected lnRR (Lajeunesse 2015) — formula and R code on the handout.
  • Never add an arbitrary constant (+1) without testing its influence.

Trap 2 · Missing or disguised SDs

The common case: a standard error SD = SE × √n
553 of our 888 comparisons were reported with a SE.

Also possible: from a 95% CI, or from a median and IQR (Wan et al. 2014) — see the formula sheet.

215 comparisons (24%) have no usable SD.
Never set it to 0 or drop silently. Robust fix for lnRR: compute the sampling variance from the average squared coefficient of variation (CV = SD/mean, weighted by n) of the other studies (Nakagawa et al. 2023). Not possible for zero means.

Traps 1 and 2 together: 224 of 888 comparisons (25%) are unusable as such for lnRR. This afternoon we set them aside — and say so.

Trap 3 · lnRR and g: same direction, different magnitudes

The 664 usable comparisons, computed both ways:

  • Same direction always (both follow X‾T−X‾C\bar X_T - \bar X_C)
  • Similar ranking (Spearman ρ = 0.84)
  • But very different magnitudes: g explodes when the SD is tiny

g depends on the SD (trap size, plot size, subsamples); lnRR does not. Never mix both in one analysis, and do not convert lnRR ↔︎ g.

Other kinds of data, in one slide

Yes/no outcomes Plants infested or not, seedlings alive or dead → a 2×2 table of counts → log odds ratio (or log risk ratio, the log of the ratio of proportions).
escalc("OR", ai=, bi=, ci=, di=)

Correlations e.g. habitat cover vs pest control → Fisher’s z.
Pool z, not r, then back-transform.
escalc("ZCOR", ri=, ni=)

Conversions between metrics Possible (log OR → d, r → d), but only for a minority of studies: flag them and run a sensitivity analysis without them. Never lnRR ↔︎ g. Formulas in the formula sheet (Borenstein et al. 2021 chap. 7).

Quick quiz

1. Earthworm biomass (g/m²) under conservation vs conventional tillage, studies in many different units. Metric?

lnRR — positive, proportional, unit-free.

2. lnRR = −0.36. Change in %?

100 × (e−0.36 − 1) = −30%.

3. Proportion of fields where a weed is present, with vs without cover crop. Metric?

log odds ratio (or log risk ratio).

4. Why use g rather than d?

d overestimates |effect| in small samples; g corrects it.

Going further

All formulas (with the exact J, bias-corrected lnRR and its variance, conversions and their variances): see the A4 formula sheet handout.

  • Variability as an outcome: lnCVR, the log ratio of coefficients of variation (Nakagawa et al. 2015).
  • Paired / before–after designs (SMCC, ROMC in escalc) need the correlation between measurements.
  • Several treatments vs one control create dependent effect sizes → multilevel models, this afternoon.

In the online notebook

Each point of this session developed further, with the same data, more pitfalls and exercises with solutions:

literaturesynthesis.github.io/notebook

References

Borenstein, M., L. V. Hedges, J. P. T. Higgins, and H. R. Rothstein. 2021. Introduction to meta-analysis. Second edition. Wiley, Chichester.
Fernandes, F. S., F. S. Ramalho, W. A. C. Godoy, J. K. S. Pachu, R. B. Nascimento, J. B. Malaquias, and J. C. Zanuncio. 2013. Within plant distribution and dynamics of hyadaphis foeniculi (Hemiptera: Aphididae) in field fennel intercropped with naturally colored cotton. Florida Entomologist 96.
Hedges, L. V. 1981. Distribution theory for Glass’s estimator of effect size and related estimators. Journal of Educational Statistics 6:107–128.
Hedges, L. V., J. Gurevitch, and P. S. Curtis. 1999. The meta-analysis of response ratios in experimental ecology. Ecology 80:1150–1156.
Jones, S. K., A. C. Sánchez, S. D. Juventia, and N. Estrada-Carmona. 2021. A global database of diversified farming effects on biodiversity and yield. Scientific Data 8:212.
Lajeunesse, M. J. 2015. Bias and correction for the log response ratio in ecological meta-analysis. Ecology 96:2056–2063.
Nakagawa, S., and I. C. Cuthill. 2007. Effect size, confidence interval and statistical significance: A practical guide for biologists. Biological Reviews 82:591–605.
Nakagawa, S., D. W. A. Noble, M. Lagisz, R. Spake, W. Viechtbauer, and A. M. Senior. 2023. A robust and readily implementable method for the meta-analysis of response ratios with and without missing standard deviations. Ecology Letters 26:232–244.
Nakagawa, S., R. Poulin, K. Mengersen, K. Reinhold, L. Engqvist, M. Lagisz, and A. M. Senior. 2015. Meta-analysis of variation: Ecological and evolutionary applications and beyond. Methods in Ecology and Evolution 6:143–152.
Wan, X., W. Wang, J. Liu, and T. Tong. 2014. Estimating the sample mean and standard deviation from the sample size, median, range and/or interquartile range. BMC Medical Research Methodology 14:135.