Learning objectives

By the end of this session, you will be able to:

  1. List the ways a set of studies can be a biased sample of the evidence.
  2. Read a contour-enhanced funnel plot — and explain why asymmetry ≠ publication bias.
  3. Test for small-study effects and time-lag bias with dependent effect sizes, and know which classic methods to avoid.
  4. Run sensitivity analyses (leave-one-study-out, influence, study quality) and report robustness.

Are our 28 studies a fair sample?

A meta-analysis assumes the studies it contains are representative of all studies done.

But studies with significant results are more likely to be:

  • published → publication bias
  • published quickly → time-lag bias
  • published in English → language bias
  • published more than once → multiple-publication bias
  • cited (and found) → citation bias

The consequence The pooled effect can be overestimated — sometimes an effect appears where there is none.

The best cure is prevention Search grey literature, theses, reports and non-English studies (Monday’s sessions). The tests we see today can only detect and probe bias — never repair it.

Adapted from Ouédraogo (2024); see also Koricheva et al. (2013), chap. 14.

The logic of the funnel plot

Large studies (top) cluster near the true effect; small ones (bottom) scatter symmetrically → a funnel.

Where would the missing studies be?
Small non-significant studies (between the dotted lines) stay in drawers → a corner is missing.

Our funnel, with significance contours

  • Shaded bands: significance of each comparison (grey background: p < 0.01)
  • Dashed line: the multilevel mean
  • Top: 83 comparisons from 2 studies with SE < 0.02 — suspiciously precise: check the extraction

If missing studies were in the white zone (non-significant), publication bias is a plausible cause. If they would sit in significant zones, look for other causes (Peters et al. 2008).

Asymmetry ≠ publication bias

Your call

Our funnel is asymmetric. List other reasons than publication bias.

  • Heterogeneity: small studies differ (other crops, pests, methods) — true small-study effects.
  • Artefacts: the variance of lnRR depends on the mean; extraction errors; pseudo-replication.
  • Chance: with few studies, funnels are often asymmetric.

Testing for small-study effects: the classic way

Egger’s test (Egger et al. 1997): regress each effect size on its standard error. A slope ≠ 0 means small studies give different results.

Our pests: p = 0.00037.

Two problems

  1. It treats our 382 comparisons as independent — they come from 28 studies.
  2. For lnRR, the SE is computed from the means themselves → an artificial link between effect and SE, even without any bias.

Testing for small-study effects: the multilevel way

Add a precision measure as a moderator of the multilevel model (Nakagawa et al. 2022), based on sample size only:

1/nT+1/nC\sqrt{1/n_T + 1/n_C}

rma.mv(yi, vi,
       mods = ~ sqrt(1/n_t + 1/n_c),
       random = ~ 1 | study_id / es_id,
       test = "t", dfs = "contain",
       data = dat)
Precision measure Slope p
SE (artefact-prone) 0.018
1/nT+1/nC\sqrt{1/n_T + 1/n_C} 0.10
same, cluster-robust (CR2) 0.18

With the recommended measure: no robust evidence of a small-study effect.

If there were bias, would the conclusion change?

Estimate Pest abundance [95% CI]
Main model -27% [-40; -11]
PET: effect of an infinitely large study -43% [-61; -18]
PEESE -36% [-50; -16]

Rule (Stanley and Doucouliagos 2014): if the PET intercept is significant (here yes), report PEESE.

Adjusted estimates are, if anything, stronger: the pest reduction is not an artefact of publication bias.

With strong heterogeneity, PET-PEESE tends to over-correct: treat it as a sensitivity check, not as “the true effect”. Alternative: selection models [Vevea and Hedges (1995); selmodel() in metafor], which model the probability of publication given the p-value.

Time-lag bias

Do early studies report stronger effects than later ones?

Year as a moderator in the multilevel model: slope = -0.0029 per year (p = 0.87).

No evidence of a decline effect — with only 28 studies, the test has little power.

Also possible: include year and SE together (Nakagawa et al. 2022).

Two classics to avoid (or to use with care)

Fail-safe N (Rosenthal) “How many null studies would make the result non-significant?”

Our data: 4,098,124 studies.

Meaningless: assumes missing studies have zero effect, ignores heterogeneity and dependence, estimates no bias. Do not report it.

Trim-and-fill Imputes “missing” studies to make the funnel symmetric.

Assumes asymmetry = publication bias, performs badly with heterogeneity, ignores dependence.
→ at most a sensitivity analysis on study-level data, never “the corrected effect”.

Source: Nakagawa et al. (2022); Ouédraogo (2024).

Does one study drive the result?

Leave-one-study-out: refit the model 28 times, dropping one whole study each time.

  • Mean ranges from -29% to -23%
  • The CI never crosses 0

Cook’s distance by study: largest = 0.32, no study stands out from the others.

The conclusion does not depend on any single study.

A robustness table for your paper

Analysis Pest abundance Conclusion changes?
Main multilevel model (28 studies) -27% [-40; -11] —
Only comparisons meeting all validity criteria† (13 studies; 130 comparisons not appraisable) -27% [-39; -13]; validity as moderator p = 0.74 No
Leave-one-study-out (range of means) -29% to -23% No
Small-study effect (sample-size based) slope p = 0.10; PEESE -36% No
Time-lag (year as moderator) slope ≈ 0 (p = 0.87) No (low power)

† Critical appraisal coded by Jones et al. (2021): enough sampling units in both groups, plots in their state long enough, intervention and comparator < 1.1 km apart. This is internal validity of each study (risk of bias, cf. CEE critical appraisal) — a different question from publication bias.

Also worth testing: another effect-size metric, the imputation of missing SDs, the correlation ρ assumed within studies. Report all sensitivity analyses — including those that change the result.

Practical · 33 min + debrief

Time Task
15:22 Contour-enhanced funnel plot (code given: interpret)
15:27 Classic vs multilevel Egger (the lnRR trap); PET-PEESE
15:37 Sensitivity: leave-one-study-out, validity
15:47 Your turn: natural enemies — or the optional parts (time-lag, Cook, fail-safe N)
15:55 Debrief (all together)

Open meta_analysis_course.Rproj (the course kit), then td_bias.R. Instructions and folded solutions: the practical web page.

In the online notebook

Each point of this session developed further, with the same data, more pitfalls and exercises with solutions:

literaturesynthesis.github.io/notebook

References

Egger, M., G. Davey Smith, M. Schneider, and C. Minder. 1997. Bias in meta-analysis detected by a simple, graphical test. BMJ 315:629–634.
Jones, S. K., A. C. Sánchez, S. D. Juventia, and N. Estrada-Carmona. 2021. A global database of diversified farming effects on biodiversity and yield. Scientific Data 8:212.
Koricheva, J., J. Gurevitch, and K. Mengersen, editors. 2013. Handbook of meta-analysis in ecology and evolution. Princeton University Press, Princeton.
Nakagawa, S., M. Lagisz, M. D. Jennions, J. Koricheva, D. W. A. Noble, T. H. Parker, A. Sánchez-Tójar, Y. Yang, and R. E. O’Dea. 2022. Methods for testing publication bias in ecological and evolutionary meta-analyses. Methods in Ecology and Evolution 13:4–21.
Ouédraogo, D.-Y. 2024. Risks of bias in meta-analyses. Lecture, FRB-CESAB training course on meta-analyses and systematic reviews.
Peters, J. L., A. J. Sutton, D. R. Jones, K. R. Abrams, and L. Rushton. 2008. Contour-enhanced meta-analysis funnel plots help distinguish publication bias from other causes of asymmetry. Journal of Clinical Epidemiology 61:991–996.
Stanley, T. D., and H. Doucouliagos. 2014. Meta-regression approximations to reduce publication selection bias. Research Synthesis Methods 5:60–78.
Vevea, J. L., and L. V. Hedges. 1995. A general linear model for estimating effect size in the presence of publication bias. Psychometrika 60:419–435.
Viechtbauer, W., and M. W.-L. Cheung. 2010. Outlier and influence diagnostics for meta-analysis. Research Synthesis Methods 1:112–125.