IIIT-H study caution doctors against relying entirely on AI for medical diagnosis
Account subscription benefits alongside Premium Stories, Editorials, Opinions and more. Unlock these with Subscription
Institute of Information Technology, Hyderabad (IIIT-H) team evaluated four models — MAIRA-2, MedGemma-4B, LLaVA-Med-1.5, and LLaVA-1.5, against thousands of publicly available chest X-rays. The image is used for representative purposes only. | Photo Credit: NISSAR AHMAD
Doctors should exercise caution when using artificial intelligence (AI) tools to interpret medical images, as areas highlighted by AI models may not always align with those identified by radiologists , according to a study by researchers at the International Institute of Information Technology, Hyderabad (IIIT-H).
The study, conducted by the Language Technologies Research Centre (LTRC) at IIIT-H, examined four vision-language models for analysing chest X-rays and found discrepancies between the regions highlighted by the models and those identified by radiologists.
The findings underscore the need for doctors to independently assess AI-generated outputs rather than relying on them entirely for diagnosis.
The research team, led by Parameswari Krishnamurthy, sought to determine whether AI-generated heatmaps accurately represented the regions of an image that radiologists would identify as disease-related.
However, the researchers questioned whether these visual indicators necessarily reflected the models’ ability to identify the actual location of a disease.
“We essentially wanted to examine whether the heatmaps created by vision-language models actually correspond to where radiologists, who look at the image, would say the disease lies,” said Syed Faizan, principal investigator of the study titled “How Well Do Chest X-Ray VLM Attention Overlays Match Radiologist Boxes? A Cross-Model Audit and Radiologist Reader.”
The team evaluated four models, MAIRA-2, MedGemma-4B, LLaVA-Med-1.5, and LLaVA-1.5, against thousands of publicly available chest X-rays. Two radiologists also participated in the study to assess them, allowing the researchers to compare AI-generated highlights with human assessments.
The study, accepted at the International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI) 2026, is to be presented at the iMIMIC Satellite Event in Strasbourg.
The researchers found that an AI model highlighting an apparently correct region of an X-ray did not necessarily mean that it had identified the disease in the same way as a radiologist.
According to Dr. Faizan, a model could arrive at a diagnosis first and then use that conclusion to determine where to place its heatmap, rather than independently identifying the affected region from the image.
To test this possibility, the team removed diagnostic information and examined how the models localised abnormalities. Their performance declined, suggesting that anatomical expectations associated with a diagnosis could influence the regions highlighted by the models.
The study also found a difference in the performance of these different models in the technical audit and radiologists’ assessments.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.thehindu.com — the content belongs to The Hindu - Sci-Tech.