Making evidence synthesis easy, the hard way
What are the prospects of using Artificial Intelligence (AI) to automate steps of the evidence synthesis pipeline? Prof Malcolm Macleod explains why there is no easy way to make evidence synthesis easy.
Prof Malcolm Macleod founded the CAMARADES group in 2005. He has been doing systematic reviews of animal studies relating to human health since before then. His group established the freely available SyRF platform to support systematic reviews, and in 2025 he was co-chair of the Evidence Synthesis Infrastructure Collaboration Working Group on the Safe and Responsible Use of AI in evidence synthesis. Recently, his group have been working with SEBI-Livestock to increase the use of automation in evidence synthesis projects.
The findings of a single study can only ever be provisional – it is through repetition, in the same and in different contexts, that we learn how broadly applicable the findings are, and how those findings should influence our decisions as research users.
Synthesising evidence across studies requires that we consider all relevant information, in ways which consider not just the research claims of contributing studies but also the provenance of those claims. The “Evidence Synthesis” community have developed sophisticated methodologies do all of this, but considerable human effort is required.
That effort, and the time taken, scales with the breadth of the research question being considered, so evidence synthesis projects tend to have a narrow scope addressing a tightly focussed research question. Where we can, reliably, substitute automated systems for human effort we could ask broader questions, consider information across research siloes, and even provide regularly updated summaries which give a real-time picture of what we know.
This promise of pain-free evidence synthesis has encouraged academic researchers and for-profit companies to develop automated approaches to some (with individual ‘digital evidence synthesis tools’ DESTs) or all (in evidence synthesis platforms) of the tasks which previously required substantial human effort. How has that gone?
In a recent Evidence Synthesis Infrastructure Collaboration Working Group on the Safe and Responsible Use of AI in evidence synthesis we carried out a ‘readiness assessment’ of AI DESTs for different stages of the review process and – in contrast to the hype that a new age had dawned – it was clear to us that for the vast majority of tasks the performance of AI DESTs remained insufficient. This has no prevented some commercial providers from making very bold claims about performance, claims which do not stand up to independent validation. But do DESTs have a place?
Comparing performance of AI-powered digital evidence synthesis tools with human-powered research
It is important, when considering their performance, not to judge this against some absolute ground truth, but to compare performance with that of the appropriately trained human for which the DEST might substitute. For citation screening, human screening in duplicate probably detects around 95% of relevant records, so this should be out target for a citation screening DEST.
Retrieval: Being able to search databases programmatically (using an Application Programming Interface, or API) allows me to download search returns in bulk, update searches on the fly, and saves hours of human time. Because it’s coded, it’s much more reproducible. It’s not ‘AI’, but it has huge impact.
Deduplication: Searching across multiple sources inevitably gives overlap in search returns, and for large corpora humans simply don’t have the working memory to process this. Fortunately, several effective deduplication DESTs are available,some of which are free. Again, it’s not AI, but has huge impact. This was covered by a recent comparative study on "Evaluating the accuracy and speed of eight deduplication tools".
Citation screening: This is one of the most time-consuming parts of the review process, and provides a major limitation to the size of the corpora that can be processed in a reasonable time (and therefore to the scope of the evidence synthesis project). The ‘standard’ current approaches generally use a regression-based analysis of word vector embeddings – each word, or group of words, in the bag of words that make up the title and abstract has meaning which can be represented in a vector, and each paper has a sum of those vectors in multidimensional space, and – through exposure to papers with human annotations – a support vector machine can ‘learn’ to distinguish between likely relevant and likely irrelevant studies. This is again not AI, but has allowed us to consider much larger corpora than we could 15 years ago. The downside is that high recall comes at a cost of lower specificity (the proportion of irrelevant studies included in error); and in our hands you need human screening decisions for training on at least 2000 records, sometimes more, before DEST performance is acceptable.
Here, there is a real prospect that AI – in this case the use of Large Language Models, will make a difference. The LLM – which may sit on a local computer, be accessed through institutional portals (such as 'ELM' at the University of Edinburgh) or directly on a provider site – is provided with a role; the purpose and inclusion criteria for the review; and is asked to determine whether the study should be included or not. While these systems don’t need ‘training’, performance does seem to depend critically on the precise wording of the prompt, and so we spend some time, using a set of human curated records, to refine prompts to maximise performance, usually asking the LLM to give a judgement on each inclusion or exclusion criteria as well as an overall judgement, so we can see where problems might lie. After this, we absolutely need to validate performance in a set of randomly selected records which have not been used for prompt refinement, to ensure that the performance generalises. It seems that different LLMs make different mistakes, so we are exploring whether combing the scores from different LLMs – or even using a second LLM to improve the prompt offered to the first – might be helpful.
Annotation and data extraction: Approaches using earlier language models like BERT in a natural language processing environment have shown reasonable performance for instance in some risk of bias annotation, and we have some work exploring LLMs for these tasks. However, for many annotation fields (the species and sex of experimental animals, the drugs used and the outcomes measured), simple regular expressions (text mining) are just as good if not better. In time, it is likely that, with careful prompt engineering, LLM based approaches applied to full text, will be transformative here as they will be for citation screening.
All of these approaches require access to article full text, and sometimes it seems that scientific publishers do everything they can to prevent users from doing this in a straightforward way. Since the authors of the original work raised the funds to do it, did the work, and paid to have it published (as well as providing, pro bono, peer review for the works of others) this seems to me to be somewhat immoral. See our blog story Does the Journal Article have a future? for more detailed discussion.
Information synthesis and meta-analysis: Providing searchable summaries of a literature – the animals studies, drugs tested and so on – is now trivially easy using the DESTs described above, and forms the backbone of our SOLES approach. Extracting quantitative outcome data, and then processing this through an appropriate meta-analysis approach, is still some way off. Where several projects are conducted on the same platform and share the same data structure it is however possible to implement an analysis pipeline (for instance in R), which usually requires only minor modification form project to project.
There are exciting developments in the automation of steps of the evidence synthesis pipeline, but it is critically important that we adopt new technologies with caution, and take care to validate their performance, so we can be confident that they perform, in our context, as advertised. There is no easy way to make evidence synthesis easy.
Header photo by Zoshua Colah on Unsplash