Publication bias, robustness and sensitivity
CIRAD, UPR HortSys
By the end of this session, you will be able to:
A meta-analysis assumes the studies it contains are representative of all studies done.
But studies with significant results are more likely to be:
The consequence The pooled effect can be overestimated — sometimes an effect appears where there is none.
The best cure is prevention Search grey literature, theses, reports and non-English studies (Monday’s sessions). The tests we see today can only detect and probe bias — never repair it.
Large studies (top) cluster near the true effect; small ones (bottom) scatter symmetrically → a funnel.
Where would the missing studies be?
Small non-significant studies (between the dotted lines) stay in drawers → a corner is missing.
If missing studies were in the white zone (non-significant), publication bias is a plausible cause. If they would sit in significant zones, look for other causes (Peters et al. 2008).
Your call
Our funnel is asymmetric. List other reasons than publication bias.
Egger’s test (Egger et al. 1997): regress each effect size on its standard error. A slope ≠ 0 means small studies give different results.
Our pests: p = 0.00037.
Two problems
Add a precision measure as a moderator of the multilevel model (Nakagawa et al. 2022), based on sample size only:
| Precision measure | Slope p |
|---|---|
| SE (artefact-prone) | 0.018 |
| 0.10 | |
| same, cluster-robust (CR2) | 0.18 |
With the recommended measure: no robust evidence of a small-study effect.
| Estimate | Pest abundance [95% CI] |
|---|---|
| Main model | -27% [-40; -11] |
| PET: effect of an infinitely large study | -43% [-61; -18] |
| PEESE | -36% [-50; -16] |
Rule (Stanley and Doucouliagos 2014): if the PET intercept is significant (here yes), report PEESE.
Adjusted estimates are, if anything, stronger: the pest reduction is not an artefact of publication bias.
With strong heterogeneity, PET-PEESE tends to over-correct: treat it as a sensitivity check, not as “the true effect”. Alternative: selection models [Vevea and Hedges (1995); selmodel() in metafor], which model the probability of publication given the p-value.
Do early studies report stronger effects than later ones?
Year as a moderator in the multilevel model: slope = -0.0029 per year (p = 0.87).
No evidence of a decline effect — with only 28 studies, the test has little power.
Also possible: include year and SE together (Nakagawa et al. 2022).
Fail-safe N (Rosenthal) “How many null studies would make the result non-significant?”
Our data: 4,098,124 studies.
Meaningless: assumes missing studies have zero effect, ignores heterogeneity and dependence, estimates no bias. Do not report it.
Trim-and-fill Imputes “missing” studies to make the funnel symmetric.
Assumes asymmetry = publication bias, performs badly with heterogeneity, ignores dependence.
→ at most a sensitivity analysis on study-level data, never “the corrected effect”.
Leave-one-study-out: refit the model 28 times, dropping one whole study each time.
Cook’s distance by study: largest = 0.32, no study stands out from the others.
The conclusion does not depend on any single study.
| Analysis | Pest abundance | Conclusion changes? |
|---|---|---|
| Main multilevel model (28 studies) | -27% [-40; -11] | — |
| Only comparisons meeting all validity criteria† (13 studies; 130 comparisons not appraisable) | -27% [-39; -13]; validity as moderator p = 0.74 | No |
| Leave-one-study-out (range of means) | -29% to -23% | No |
| Small-study effect (sample-size based) | slope p = 0.10; PEESE -36% | No |
| Time-lag (year as moderator) | slope ≈ 0 (p = 0.87) | No (low power) |
† Critical appraisal coded by Jones et al. (2021): enough sampling units in both groups, plots in their state long enough, intervention and comparator < 1.1 km apart. This is internal validity of each study (risk of bias, cf. CEE critical appraisal) — a different question from publication bias.
Also worth testing: another effect-size metric, the imputation of missing SDs, the correlation ρ assumed within studies. Report all sensitivity analyses — including those that change the result.
| Time | Task |
|---|---|
| 15:22 | Contour-enhanced funnel plot (code given: interpret) |
| 15:27 | Classic vs multilevel Egger (the lnRR trap); PET-PEESE |
| 15:37 | Sensitivity: leave-one-study-out, validity |
| 15:47 | Your turn: natural enemies — or the optional parts (time-lag, Cook, fail-safe N) |
| 15:55 | Debrief (all together) |
Open meta_analysis_course.Rproj (the course kit), then td_bias.R. Instructions and folded solutions: the practical web page.
Each point of this session developed further, with the same data, more pitfalls and exercises with solutions:
literaturesynthesis.github.io/notebook