Skip to main content
header image

Coker Babafemi

Coker is a medical doctor with a developing specialisation in health data science. He is currently studying for an MRes in Data Science for Health at the University of Liverpool.

Pilot for searching NHS imaging records


Project Overview

Hospitals hold millions of free-text radiology reports, but few are labelled. This project tested how reliably reports can be labelled automatically, to support training-data creation and clinical triage. Using 57,805 CheXpert Plus chest X-ray reports, I compared keyword search, RadGraph-based rules and ten large language models (LLMs), benchmarked against 149 hand-labelled reports. A second aim was to measure how much simple word search misses when a pathology is implied rather than named. For this, I built a search app comparing keyword matches with LLM-generated differential diagnoses.

What were the key results of your research project?

Keyword searches fails on negation. Radiologists routinely report what is absent ("no pleural effusion"), so word matches cannot separate normal from abnormal findings. Embedding-based semantic search shares this weakness.

RadGraph helps with these issues but has its limits. A rule combining measured and present observations isolated 2,119 difficult reports (3.7%), but case review showed missing relations, no sense of time, and normal findings also tagged "present".

LLMs performed strongly when addressing this problem, where with a findings-extraction prompt, every model caught at least 97% of abnormal reports it answered. Claude 3.7 Sonnet led with 99.3% recall, missing one of 143 abnormal reports.

Errors in results reflected definitions, not mis-readings where remaining disagreements involved benign calcified granulomas and misplaced lines, so hinged on how "abnormal" is defined.

The LLMs surface implied pathology. Searching "malignancy" across the 149 reports found no matches in the report text, 3 with synonyms, and 53 when LLM differentials were included.

The limitations of the analysis was that only six normal reports were in the test set, so specificity needs a larger labelled sample.

How did you improve your technical skills?

  • Clinical text is hard to parse. I learned first-hand how negation, past findings and varying reporting styles make medical records difficult to interpret automatically.
  • RadGraph: I learned how it represents reports as structured graphs, and where it falls short.
  • Embeddings: I came to understand vector embeddings and why semantic search alone struggles with clinical meaning.
  • High-Performance Computing: I accessed and used an HPC cluster to run large language models.
  • LLMs through APIs, not chat: I moved beyond chat interfaces to working with LLMs programmatically. That meant designing prompts that reliably produce the output I need, handling and validating model responses, and seeing how settings such as batch size and temperature affect results.


GitHub repository: https://github.com/Doctorstrange/cir




How do you feel you have benefitted from completing this internship and has it made you consider future career paths?

This internship gave me a realistic view of what it takes to plan and run a research project from start to finish: defining the question, preparing the data, testing ideas, and revising my approach when the results were unexpected.

The experience has shaped my career direction. I want to work on fine-tuning and developing AI models for healthcare. So much of clinical practice happens through language: notes, reports, referrals and handovers. Because current models understand the world mainly through language, making them more reliable on clinical text could bring large efficiency gains for clinicians and better outcomes for patients. I'm particularly interested in exploring the limits of language-based AI in medicine, especially as newer approaches such as world models emerge, and in helping build systems that clinicians can trust.



Return to article index