Common-effect, random-effects and multilevel models — and how to show the results
CIRAD, UPR HortSys
By the end of this session, you will be able to:
metafor (practical, 33 min).This morning, escalc() gave us, for each of the 382 comparisons of pest abundance:
yi: the lnRR (effect size)vi: its sampling varianceThese come from 28 studies.
This afternoon’s question How do we turn 382 numbers into one answer — and an honest statement of how certain and how variable it is?
Plain average of the 382 lnRR: -24%. Why is that a bad answer? (think 20 s)
— studies differ only by sampling error.
Also called “fixed-effect” or “equal-effects” (method = "EE").
, with — true effects really differ (soils, climates, pests…).
Common effect:
Random effects:
Adding τ² to every variance makes the weights more equal: very precise studies dominate less.
| Share of the total weight | Common effect | Random effects |
|---|---|---|
| Zhou 2013 (51 very precise rows) | 76% | 17% |
| Maluleke 2005 (80 rows) | 23% | 26% |
Common effect: 2 studies out of 28 carry 99% of the weight.
Random effects: τ² tames the precise studies — but a study with many rows still counts a lot. That is the next problem.
Show of hands
Q is significant (p < 0.001). Should I switch from a common-effect to a random-effects model?
No — the choice is conceptual, not a test result. Do you want to generalise beyond these exact studies? Do you expect true effects to differ (sites, species, practices)? In ecology and agronomy: yes and yes → random effects, stated in the protocol. The Q test has low power with few studies and is almost always significant with many.
| Meaning | Our pests (random-effects model, dependence ignored for now) | |
|---|---|---|
| τ² (τ) | Variance (SD) of the true effects, in lnRR units | τ = 0.58 |
| I² | Share of the observed variance not due to sampling error — a relative measure | 99.9% |
| Q | Test of “all true effects are equal” | p < 0.001 |
I² ≈ 100% does not mean “huge differences”: it means sampling error is tiny compared with them — typical when studies are large. I² says nothing about how different the effects are. For that, use τ and the prediction interval.
CI: -33% to -22% — how precisely we know the average. Shrinks as studies accumulate.
PI: -77% to +128% — where the effect in a new field is expected. Reflects τ²; does not shrink (IntHout et al. 2016).
1. I² = 0%. Are the true effects identical?
Not necessarily: with few, imprecise studies, heterogeneity is simply not detectable.
2. A farmer asks: “will intercropping reduce pests on my farm?” CI or PI?
PI — and here it includes increases.
3. τ = 0.58 on the lnRR scale. Small or large?
Large: mean ± 1 τ already spans -60% to +29%.
Several taxa, dates, sites or treatments within the same study: same team, same field, same weather, sometimes the same control plots.
Treating them as independent = pseudo-replication: CIs too narrow, studies with many comparisons over-weighted.
study_id / es_id = rows nested in studies:
Our pests Of the total variance, 52% is between studies and 48% within studies (multilevel I², Nakagawa and Santos (2012)):
ṽ = a “typical” sampling variance.
test = "t", dfs = "contain": t-based CIs whose degrees of freedom follow the number of studies (28 − 1 = 27), not of comparisons (381) — safer than z.
Show of hands
Same 382 comparisons, four models. Which one gives the widest confidence interval — and does the mean change much?
The first two models are overconfident: one study dominates (common effect) or 382 comparisons are treated as 382 studies.
Robust variance estimation (robust(), clustered by study) is a safety net when the dependence structure is uncertain (Pustejovsky and Tipton 2022). Needs ≳ 10–20 studies.
yi), sized by precisionci.lb, ci.ubpi.lb, pi.ubpred (all from predict(model))With hundreds of effect sizes, a classic forest plot is unreadable; the orchard plot shows mean, uncertainty, heterogeneity and raw data at once (Nakagawa et al. 2021).
| Time | Task |
|---|---|
| 14:22 | Common vs random effects: estimates and weights |
| 14:30 | Heterogeneity: τ², I², prediction interval |
| 14:38 | Multilevel model + robust() |
| 14:46 | Orchard plot, then your turn: change one word, rerun for natural enemies |
| 14:55 | Debrief (all together) |
Open meta_analysis_course.Rproj (the course kit), then td_models.R. Instructions and folded solutions: the practical web page.
Each point of this session developed further, with the same data, more pitfalls and exercises with solutions:
literaturesynthesis.github.io/notebook