Meta-regression, missing data and a glimpse of network meta-analysis
CIRAD, UPR HortSys
By the end of this session, you will be able to:
Pests under intercropping: -27% on average…
…but in a new field anywhere from -80% to +170%.
Why? Crop, pest species, climate, design, intercrop density… A moderator is a study or comparison characteristic that may explain part of this variation.
At 15:00 we already did it The multilevel Egger test added precision as a moderator; the validity test added study quality. Meta-regression = a meta-analysis with predictors.
This afternoon’s multilevel model, plus a predictor:
Three things to read
rma.mv(yi, vi,
mods = ~ functional_group,
random = ~ 1 | study_id / es_id,
test = "t", dfs = "contain", data = dat)| In the equation | In the code | Meaning |
|---|---|---|
| β1 xij | mods = ~ functional_group |
the moderator |
| uj | study_id |
between-study heterogeneity |
| uij | es_id |
within-study heterogeneity |
| εij | vi |
sampling error (known) |
Pseudo-R² = share of the total heterogeneity removed by the moderator, computed on the same rows:
A clear moderator… that explains only ~10% of the heterogeneity. Most of the variation remains unexplained.
Your call
Studies on pests and studies on natural enemies may differ in crops, regions, designs… Can we trust the difference?
Between-study differences are observational: any study characteristic that travels with the moderator can confound it.
14 studies measured both groups. Split the moderator into its within- and between-study parts (mods = ~ pest_w + pest_m):
within studies: -36% · between studies: -28% (p = 0.09).
The within-study contrast — same fields, same year — is the stronger evidence.
| Rule | Why |
|---|---|
| Choose moderators a priori, in the protocol | Testing many moderators guarantees false positives |
| ≈ 10 studies (not rows) per moderator level or coefficient | A between-study moderator is tested on the number of studies (remember df ≈ 10 above): many rows per study add little |
| Report the omnibus test, coefficients and residual heterogeneity | “Significant” is not “explains the variation” |
| Beware confounding between moderators | Crops, regions and designs travel together |
| Prefer within-study contrasts when they exist | They control for study context |
| Moderator effects are associations, not causes | Studies were not randomised to moderator levels |
To impute a missing SD for lnRR, we borrow the typical coefficient of variation of the other studies (Nakagawa et al. 2023a).
First, look at the CVs: 2 studies stand out (CV > 3), one with n = 1,200 per group. Their SDs were rebuilt as SE × √n: if n counts subsamples rather than true replicates, both the SD and the precision are wrong. Check the paper.
Pests (suspect studies set aside): 53 of 376 comparisons have no usable SD — 45 of them from one study.
Why the CV? The variance of lnRR is
So a missing SD only needs a typical CV, borrowed from the other studies:
: mean CV weighted by n, per group [Nakagawa et al. (2023a); metafor vtype = "AV"]. The screening matters: squared mean CV = 0.21 without the suspect studies, 28 with them.
| Pest abundance [95% CI] | |
|---|---|
| Complete cases (323) | -28% [-42; -11] |
| With imputed SDs (376) | -30% [-45; -11] |
Same conclusion — but most imputed rows come from one study, so do not over-generalise. Report both. Effect sizes: recompute them (from t, p, figures, medians), never fill them with averages.
Intercropping or agroforestry: which is better? Only 1 of 237 studies tested both.
Network meta-analysis combines direct and indirect evidence (A vs C and B vs C → A vs B) to rank practices (Salanti 2012).
Here: a star around Monoculture, no closed loop → consistency cannot be checked; and transitivity is doubtful (the “monoculture” of agroforestry studies ≠ that of intercropping studies). Tools: netmeta, rma.mv.
| Method | When | Keep in mind |
|---|---|---|
| Multivariate meta-analysis | Several correlated outcomes (e.g. biodiversity and yield on the same plots) | Needs the within-study correlation — test a range (van Houwelingen et al. 2002) |
| Selection models | Sensitivity to publication bias | Need many studies; ignore dependence |
| Random forests (MetaForest) | Exploring many candidate moderators | Exploratory only; weight by precision; cross-validate by study |
| Bayesian meta-analysis | Few studies, informative priors, complex structures | Priors must be justified and tested |
| Time | Task | You… |
|---|---|---|
| 16:33 | Meta-regression: pests vs natural enemies | fill 1 blank, read F, coefficients, R² |
| 16:40 | Within vs between studies | fill 1 blank, compare the two estimates |
| 16:47 | Missing SDs: CVs, impute, compare | run the given code, interpret |
| 16:55 | Debrief and wrap-up of the day (until 17:00) |
The script is almost complete: the work is reading the outputs, not writing code. Questions to discuss are in the script (# Q:). Open meta_analysis_course.Rproj (the course kit), then td_advanced.R; folded answers on the practical web page.
Each point of this session developed further, with the same data, more pitfalls and exercises with solutions:
literaturesynthesis.github.io/notebook