Learning objectives

By the end of this session, you will be able to:

  1. Explain why studies are weighted, and how the weights change between models.
  2. Distinguish common-effect, random-effects and multilevel models — and choose one before seeing the data.
  3. Read μ, its CI, τ², I² and the prediction interval, and say which answers which question.
  4. Fit and plot them in R with metafor (practical, 33 min).

Where we are

This morning, escalc() gave us, for each of the 382 comparisons of pest abundance:

  • yi: the lnRR (effect size)
  • vi: its sampling variance

These come from 28 studies.

This afternoon’s question How do we turn 382 numbers into one answer — and an honest statement of how certain and how variable it is?

Plain average of the 382 lnRR: -24%. Why is that a bad answer? (think 20 s)

Two views of the world

yk=μ+εky_k = \mu + \varepsilon_k — studies differ only by sampling error.
Also called “fixed-effect” or “equal-effects” (method = "EE").

yk=μ+uk+εky_k = \mu + u_k + \varepsilon_k, with Var(uk)=τ2\text{Var}(u_k) = \tau^2 — true effects really differ (soils, climates, pests…).

Weights: precise studies count more

Common effect: wk=1vkw_k = \frac{1}{v_k}

Random effects: wk=1vk+τ2w_k = \frac{1}{v_k + \tau^2}

Adding τ² to every variance makes the weights more equal: very precise studies dominate less.

Share of the total weight Common effect Random effects
Zhou 2013 (51 very precise rows) 76% 17%
Maluleke 2005 (80 rows) 23% 26%

Common effect: 2 studies out of 28 carry 99% of the weight.
Random effects: τ² tames the precise studies — but a study with many rows still counts a lot. That is the next problem.

Source: Zhou et al. (2013); Maluleke et al. (2005); via Jones et al. (2021).

Which model? Decide before seeing the data

Show of hands

Q is significant (p < 0.001). Should I switch from a common-effect to a random-effects model?

No — the choice is conceptual, not a test result. Do you want to generalise beyond these exact studies? Do you expect true effects to differ (sites, species, practices)? In ecology and agronomy: yes and yes → random effects, stated in the protocol. The Q test has low power with few studies and is almost always significant with many.

Three numbers, three meanings

Meaning Our pests (random-effects model, dependence ignored for now)
τ² (τ) Variance (SD) of the true effects, in lnRR units τ = 0.58
I² Share of the observed variance not due to sampling error — a relative measure 99.9%
Q Test of “all true effects are equal” p < 0.001

I² ≈ 100% does not mean “huge differences”: it means sampling error is tiny compared with them — typical when studies are large. I² says nothing about how different the effects are. For that, use τ and the prediction interval.

Confidence vs prediction interval

CI: -33% to -22% — how precisely we know the average. Shrinks as studies accumulate.

PI: -77% to +128% — where the effect in a new field is expected. Reflects τ²; does not shrink (IntHout et al. 2016).

Quick quiz

1. I² = 0%. Are the true effects identical?

Not necessarily: with few, imprecise studies, heterogeneity is simply not detectable.

2. A farmer asks: “will intercropping reduce pests on my farm?” CI or PI?

PI — and here it includes increases.

3. τ = 0.58 on the lnRR scale. Small or large?

Large: mean ± 1 τ already spans -60% to +29%.

382 comparisons are not 382 studies

Several taxa, dates, sites or treatments within the same study: same team, same field, same weather, sometimes the same control plots.

Treating them as independent = pseudo-replication: CIs too narrow, studies with many comparisons over-weighted.

The multilevel (three-level) model

yij=μ+uj+uij+εijy_{ij} = \mu + u_{j} + u_{ij} + \varepsilon_{ij}

  • uju_{j}: between-study variation, σstudy2\sigma^2_{study}
  • uiju_{ij}: within-study variation (between comparisons), σwithin2\sigma^2_{within}
  • εij\varepsilon_{ij}: sampling error, vijv_{ij} (known)

study_id / es_id = rows nested in studies:

rma.mv(yi, vi,
       random = ~ 1 | study_id / es_id,
       test = "t", dfs = "contain",
       data = dat)

Our pests Of the total variance, 52% is between studies and 48% within studies (multilevel I², Nakagawa and Santos (2012)):

Istudy2=σstudy2σstudy2+σwithin2+ṽI^2_{study} = \frac{\sigma^2_{study}}{\sigma^2_{study} + \sigma^2_{within} + \tilde v}

ṽ = a “typical” sampling variance.

test = "t", dfs = "contain": t-based CIs whose degrees of freedom follow the number of studies (28 − 1 = 27), not of comparisons (381) — safer than z.

Predict first

Show of hands

Same 382 comparisons, four models. Which one gives the widest confidence interval — and does the mean change much?

Same data, four models

The first two models are overconfident: one study dominates (common effect) or 382 comparisons are treated as 382 studies.

Robust variance estimation (robust(), clustered by study) is a safety net when the dependence structure is uncertain (Pustejovsky and Tipton 2022). Needs ≳ 10–20 studies.

Beyond the forest plot: the orchard plot

  • Points: every comparison (yi), sized by precision
  • Thick bar: 95% CI of the mean — ci.lb, ci.ub
  • Thin bar: 95% prediction interval — pi.lb, pi.ub
  • Dot: the mean — pred (all from predict(model))
  • k and number of studies

With hundreds of effect sizes, a classic forest plot is unreadable; the orchard plot shows mean, uncertainty, heterogeneity and raw data at once (Nakagawa et al. 2021).

Practical · 33 min + debrief

Time Task
14:22 Common vs random effects: estimates and weights
14:30 Heterogeneity: τ², I², prediction interval
14:38 Multilevel model + robust()
14:46 Orchard plot, then your turn: change one word, rerun for natural enemies
14:55 Debrief (all together)

Open meta_analysis_course.Rproj (the course kit), then td_models.R. Instructions and folded solutions: the practical web page.

In the online notebook

Each point of this session developed further, with the same data, more pitfalls and exercises with solutions:

literaturesynthesis.github.io/notebook

References

Borenstein, M., L. V. Hedges, J. P. T. Higgins, and H. R. Rothstein. 2021. Introduction to meta-analysis. Second edition. Wiley, Chichester.
IntHout, J., J. P. A. Ioannidis, M. M. Rovers, and J. J. Goeman. 2016. Plea for routinely presenting prediction intervals in meta-analysis. BMJ Open 6:e010247.
Jones, S. K., A. C. Sánchez, S. D. Juventia, and N. Estrada-Carmona. 2021. A global database of diversified farming effects on biodiversity and yield. Scientific Data 8:212.
Maluleke, M. H., A. Addo-Bediako, and K. K. Ayisi. 2005. Influence of maize/lablab intercropping on lepidopterous stem borer infestation in maize. Journal of Economic Entomology 98:384–388.
Nakagawa, S., M. Lagisz, R. E. O’Dea, J. Rutkowska, Y. Yang, D. W. A. Noble, and A. M. Senior. 2021. The orchard plot: Cultivating a forest plot for use in ecology, evolution, and beyond. Research Synthesis Methods 12:4–12.
Nakagawa, S., and E. S. A. Santos. 2012. Methodological issues and advances in biological meta-analysis. Evolutionary Ecology 26:1253–1274.
Nakagawa, S., Y. Yang, E. L. Macartney, R. Spake, and M. Lagisz. 2023. Quantitative evidence synthesis: A practical guide on meta-analysis, meta-regression, and publication bias tests for environmental sciences. Environmental Evidence 12:8.
Pustejovsky, J. E., and E. Tipton. 2022. Meta-analysis with robust variance estimation: Expanding the range of working models. Prevention Science 23:425–438.
Rice, K., J. P. T. Higgins, and T. Lumley. 2018. A re-evaluation of fixed effect(s) meta-analysis. Journal of the Royal Statistical Society: Series A 181:205–227.
Senior, A. M., C. E. Grueber, T. Kamiya, M. Lagisz, K. O’Dwyer, E. S. A. Santos, and S. Nakagawa. 2016. Heterogeneity in ecological and evolutionary meta-analyses: Its magnitude and implications. Ecology 97:3293–3299.
Zhou, H.-B., J.-L. Chen, Y. Liu, F. Francis, E. Haubruge, C. Bragard, J.-R. Sun, and D.-F. Cheng. 2013. Influence of garlic intercropping or active emitted volatiles in releasers on aphid and related beneficial in wheat fields in China. Journal of Integrative Agriculture 12.