Understanding, choosing and computing the currency of meta-analysis
CIRAD, UPR HortSys
By the end of this session, you will be able to:
metafor::escalc().The 888 intercropping comparisons of our running example report biodiversity as:
An effect size puts all studies on one common scale: the direction and magnitude of the difference between intervention and comparator — plus its uncertainty.
Source: Jones et al. (2021); the authors converted everything to SD before sharing the data.
A p-value mixes effect size and sample size: a tiny effect can be “highly significant”, a large one “not significant”.
Meta-analysis combines effect sizes and their uncertainty, not p-values (Nakagawa and Cuthill 2007).
Each effect size comes with its sampling variance . In the meta-analysis this afternoon, each study is weighted by the inverse of its variance, : precise studies count more. (In a random-effects model, the weight also includes the real variation between studies, : .)
Proportional (→ lnRR): “−30% aphids”, same meaning with 10 or 1,000 aphids. Abundance, biomass, yield: a true zero, effects multiply.
Additive (→ g): “+5 units”. Scores, indices, variables with an arbitrary zero.
In ecology and agronomy, lnRR and Hedges’ g (a standardised mean difference, SMD) dominate. Our running example → lnRR. Fix the metric in the protocol, before seeing the results.
Show of hands: which metric for your own data?
Aphids on fennel, intercropped with cotton vs fennel monoculture (Fernandes et al. 2013), one comparison of our dataset:
| Intercropping (T) | Monoculture (C) | |
|---|---|---|
| Mean | 8.37 | 23.98 |
| SD | 2.85 | 5.46 |
| n | 4 | 4 |
Convention for the whole week: T (intervention) in the numerator / first, C (comparator) second. A negative effect = fewer pests with intercropping.
Back to percent for reporting:
Why the log? The ratio is asymmetric (halving = 0.5, doubling = 2); its log is symmetric (−0.69, +0.69) and closer to normal.
Why ecologists like it Unit-free and proportional: “intercropping reduces aphids by 65%” speaks to agronomists — and its value does not depend on the SD (only its variance does).
Calculator ready
With the aphid example (T: 8.37 ± 2.85, n = 4; C: 23.98 ± 5.46, n = 4), compute lnRR, its variance, and the % change.
lnRR = ln(8.37 / 23.98) = -1.053
v = 2.85² / (4 × 8.37²) + 5.46² / (4 × 23.98²) = 0.0419 (SE = 0.205)
Change = 100 × (e-1.053 − 1) = -65%
Predict first What value of yi should R return for our aphid example?
library(metafor)
# ic <- read.csv("intercropping_biodiversity.csv") # running example
dat <- escalc(measure = "ROM", # lnRR ("ratio of means")
m1i = b_mean_t, sd1i = b_sd_t, n1i = b_n_t, # T = intercropping
m2i = b_mean_c, sd2i = b_sd_c, n2i = b_n_c, # C = monoculture
data = ic)
# other metrics: measure = "SMD" (Hedges' g), "OR", "ZCOR"
dat[dat$es_id == 2011, c("es_id", "yi", "vi")] # our worked example
es_id yi vi
137 2011 -1.0526 0.0419
NA_yi vi_zero
156 46
yi = effect size, vi = sampling variance: the input of every meta-analysis (rma(), this afternoon).NA; zero SD → vi = 0 (“infinite” precision!): count both before rma().Idea: express the difference in standard-deviation units.
Useful when studies use different scales and the effect is additive.
Small samples inflate d → multiply by a correction J < 1: that is Hedges’ g (Hedges 1981).
With our aphids:
1. Difference: 8.37 − 23.98 = -15.61 aphids
2. Pooled SD (equal n: average the variances): √((2.85² + 5.46²) / 2) = 4.36
3. d = -15.61 / 4.36 = -3.58 standard deviations
4. With 4 + 4 plants, J = 0.87 → g = -3.12
escalc(measure = "SMD") does all of this — and the variance. You never compute g by hand in practice.
Your call
156 comparisons in our data have a zero mean. What do you do with them for a lnRR?
Exact zeros: no lnRR (and most also have SD = 0, so no g). Report them; test the conclusion with and without.
Near zero: lnRR is biased. Geary’s rule: each mean ≥ 3 SE from zero → 28% of our usable comparisons fail it.
What to do
The common case: a standard error SD = SE × √n
553 of our 888 comparisons were reported with a SE.
Also possible: from a 95% CI, or from a median and IQR (Wan et al. 2014) — see the formula sheet.
215 comparisons (24%) have no usable SD.
Never set it to 0 or drop silently. Robust fix for lnRR: compute the sampling variance from the average squared coefficient of variation (CV = SD/mean, weighted by n) of the other studies (Nakagawa et al. 2023). Not possible for zero means.
Traps 1 and 2 together: 224 of 888 comparisons (25%) are unusable as such for lnRR. This afternoon we set them aside — and say so.
The 664 usable comparisons, computed both ways:
g depends on the SD (trap size, plot size, subsamples); lnRR does not. Never mix both in one analysis, and do not convert lnRR ↔︎ g.
Yes/no outcomes Plants infested or not, seedlings alive or dead → a 2×2 table of counts → log odds ratio (or log risk ratio, the log of the ratio of proportions).
escalc("OR", ai=, bi=, ci=, di=)
Correlations e.g. habitat cover vs pest control → Fisher’s z.
Pool z, not r, then back-transform.
escalc("ZCOR", ri=, ni=)
Conversions between metrics Possible (log OR → d, r → d), but only for a minority of studies: flag them and run a sensitivity analysis without them. Never lnRR ↔︎ g. Formulas in the formula sheet (Borenstein et al. 2021 chap. 7).
1. Earthworm biomass (g/m²) under conservation vs conventional tillage, studies in many different units. Metric?
lnRR — positive, proportional, unit-free.
2. lnRR = −0.36. Change in %?
100 × (e−0.36 − 1) = −30%.
3. Proportion of fields where a weed is present, with vs without cover crop. Metric?
log odds ratio (or log risk ratio).
4. Why use g rather than d?
d overestimates |effect| in small samples; g corrects it.
All formulas (with the exact J, bias-corrected lnRR and its variance, conversions and their variances): see the A4 formula sheet handout.
escalc) need the correlation between measurements.Each point of this session developed further, with the same data, more pitfalls and exercises with solutions:
literaturesynthesis.github.io/notebook