Can Large Language Models Review the Scientific Literature?
Project Overview
This project tested whether AI tools can carry out a scientific literature review, benchmarked against an expert review of 102 papers on the Spanish plume, which was withheld from the systems. It measured three stages separately: finding the papers, judging how each paper uses the term, and writing a review that two experts assessed blind.
What were the key results of your research project?
- No single system found all the relevant papers. Every one missed a substantial part of the literature.
- What a system found depended mainly on access, not on the model. Running a coding agent in a browser logged in to university library subscriptions sharply increased what it retrieved.
- Coverage was uneven across journals, which points to publisher platforms rather than the AI as the limiting factor.
- All three systems correctly kept every paper that engages substantially with the term. - But all three treated many passing mentions as substantial. Their errors all ran in the same direction: making papers look more important than they are.
- Two experts, reading the reviews blind, found them readable and mostly accurate but rarely insightful: largely summaries, sometimes overstating how much the papers agree, and sometimes not answering the question asked.
- The two experts also disagreed with each other about how accurate and useful the reviews were, so judging AI-written reviews is itself an unsolved problem.
- Overall, these tools can help find and organize a literature, but the evidence and the evaluation still need expert checking.
How do you feel you have benefitted from completing this internship and has it made you consider future career paths?
I have benefited from the internship by gaining experience applying and evaluating LLMs in a real research setting, rather than only using them for general tasks. I developed a better understanding of how AI tools can support stages of the research process, such as literature searching, interpreting papers and synthesising findings, as well as where their limitations can affect the reliability of the results. The project also gave me experience comparing different models, designing evaluation methods and communicating research findings.
It has strengthened my interest in pursuing a career in machine learning and AI, particularly in research-focused roles. It has also made me more interested in postgraduate study and in working on projects where machine learning can be applied to real scientific or technical problems.