
Presentation Master's thesis - Paulo Zirlis - Psychological Methods
Presentation Master's thesis - Paulo Zirlis - Psychological Methods
- Startdatum
- 09-09-2026 13:00
- Einddatum
- 09-09-2026 14:00
- Locatie
Large language models (LLMs) have shown promise for automating parts of systematic review workflows, but their ability to recover meta-analysis-ready effect sizes from primary studies remains unclear. This study evaluated GPT-4o, GPT-5.6, and Claude Haiku-4.5 across two extraction strategies: a domain-customised strategy adapted from Li et al. (2025) and a meta-analysis-informed research-question (RQ) strategy. Effect sizes reported in 28 published meta-analyses across five research domains served as criterion-reference values, yielding 473 scoreable effects from 369 primary studies. Performance was assessed using precision, recall, F1, numerical-discrepancy measures, and mixed-effects models. Overall effect-size recovery was low: the best-performing condition, GPT-5.6 with the RQ strategy, achieved precision of .329, recall of .116, and F1 of .172 under the primary two-decimal agreement criterion.
The RQ strategy improved recovery overall, but this advantage was strongly model-dependent and concentrated in GPT-5.6. Element-level recall in the replication analysis closely matched Li et al. (2025), whereas complete effect-size recovery was substantially poorer. Multi-target studies were particularly difficult, with no correct recoveries among 937 estimable, resolved complex-study observations. Numerical agreement among successfully comparable effects was stronger than overall recovery suggested. These findings indicate that the principal bottleneck lies in identifying and aligning the information required for the intended meta-analytic effect. Current LLMs are therefore better suited to assistive workflows with human verification than to autonomous effect-size extraction.