A new artificial intelligence system aims to read the images produced during an upper endoscopy and draft the diagnosis and report on its own, according to a report published by Bioengineer.org.

The work centers on what its authors call a "bootstrapped multimodal large language model." In plain terms, that means an AI built to handle more than one kind of input — in this case, both the visual scans from a procedure and the medical language used to describe them. "Multimodal" refers to that mix of image and text, while "bootstrapped" points to a training approach in which the model builds up its own medical knowledge rather than relying solely on hand-labeled examples.

The target procedure is an EGD, short for esophagogastroduodenoscopy — the upper endoscopy exam in which a doctor threads a camera down through the esophagus, stomach, and the first part of the small intestine to look for problems. According to Bioengineer.org, the model is designed to learn the relevant medical knowledge and then perform automatic diagnosis and reporting for these exams.

The source item does not detail how accurate the system is, how it was tested, or whether it has been used in real clinical settings, so those questions remain open based on the information available.

Why it matters: endoscopy generates large volumes of images that specialists must review and write up by hand, and tools that can help interpret those scans and draft reports could ease that workload — making early, reliable evaluation of such systems important before they reach patients.