[ECE Department] Professor Jaeyoung Do’s research team selected as CVPR 2026 Award Candidate and for Oral Presentation

Professor Jaeyoung Do’s research team at the AIDAS Lab in the Department of Electrical and Computer Engineering at Seoul National University announced that its medical vision-language model (VLM), MEDIC-AD, has been accepted as an Oral Presentation and selected as an Award Candidate at CVPR 2026, one of the world’s most prestigious conferences in artificial intelligence and computer vision. MEDIC-AD was developed through clinical collaboration with Samsung Medical Center and joint research with the NVIDIA AI Technology Center (NVAITC), operated by NVIDIA, a global leader in semiconductors and AI. This research was designed to address a key limitation of existing medical AI models in that they possess broad medical knowledge but often lack the capabilities essential for real-world clinical practice, namely lesion detection, longitudinal symptom tracking, and visually explainable reasoning.
Addressing Real Challenges in Clinical Practice
MEDIC-AD focuses on solving three core tasks required in actual clinical settings: lesion detection, symptom tracking, and explainability. While most existing medical AI models have focused on acquiring vast amounts of medical knowledge, clinical practice requires more than knowledge alone. AI must be able to accurately identify abnormal lesions in medical images, determine whether a disease has improved or worsened compared with previous scans, and provide visual evidence that clinicians can verify. MEDIC-AD distinguishes itself from prior research by integrating all three capabilities into a single model.
In particular, if AI can accurately classify patient progress as “no change,” “improved,” or “worsened” during follow-up, it can reduce the diagnostic burden on clinicians and help detect subtle changes at an early stage. The model is expected to have significant clinical impact, including the detection of early-stage lesions that may otherwise be missed and the rapid assessment of treatment response.
Core AI Technology: Clinical Intelligence Built in Three Stages
The key technical feature of MEDIC-AD is its stage-wise framework, a sequential learning structure composed of three stages.
The first stage is anomaly detection. The research team inserted a new learning component, called an anomaly-aware token, into the transformer layers of the vision-language model. This token generates an Anomaly Attention Map, a probability map that distinguishes normal patches from abnormal patches, enabling the model to focus more effectively on lesion regions. Because this structure was trained across various imaging modalities, including brain MRI, head CT, and chest X-ray, the model can detect new diseases not included in the training data in a zero-shot setting.

Figure 1. Architecture of the MEDIC-AD model
The second stage is difference reasoning. Existing models typically process two images by simply concatenating them, which limits their ability to capture clinically meaningful changes over time. MEDIC-AD introduces a difference token that explicitly compares and separates anomaly features extracted from previous and current images of the same patient. This allows the model to precisely infer disease progression by identifying actual changes in lesions, without being misled by non-clinical variations such as overall brightness changes or differences in imaging angle.

Figure 2. Example of MEDIC-AD evaluated on the MMXU benchmark. The model detects changes in findings between two X-ray images and visualizes the evidence behind its assessment.
The third stage is visual explainability. To improve the reliability of AI-assisted diagnosis, the model must be able to show clinicians why it reached a particular conclusion. MEDIC-AD combines the tokens learned in the first stage with a ConvNeXt-based segmentation head to visualize, as a heatmap, the specific image regions that served as the basis for the model’s judgment. This is a key function that aligns the AI model’s conclusions with visual evidence, thereby improving clinical trust.
Global Validation Through Joint Research with NVAITC
In this research, collaboration with the NVIDIA AI Technology Center (NVAITC) went beyond computing support. It involved joint research across large-scale model optimization and overall research direction. Collaboration with NVIDIA, which possesses world-class AI infrastructure and expertise, contributed to improving the model’s robustness and global competitiveness.
On the clinical data side, the team obtained long-term follow-up chest X-ray data from 300 real patients through collaboration with Professor Pa Hong’s research team at Samsung Changwon Hospital. Unlike typical AI studies that train and evaluate models on public benchmark datasets, this study validated model performance using data collected from real hospital workflows, further strengthening its potential for clinical application.
Superior Performance Compared with Global Models Including GPT-4o and Claude
The results showed that MEDIC-AD outperformed existing medical AI models as well as leading global large language models, including OpenAI’s GPT-4o and Anthropic’s Claude 3.5, across all three tasks: lesion detection, symptom tracking, and visual explainability. In particular, on MMXU, a benchmark for disease-change analysis based on long-term clinical data, MEDIC-AD achieved an overall accuracy of 65.5%, substantially outperforming next-generation foundation models such as Lingshu at 62.0% and Citrus-V at 57.1%. In the visual explainability metric mIoU, which evaluates heatmap quality, MEDIC-AD achieved a score of up to 87.6, far exceeding competing models such as Citrus-V, which recorded an mIoU of 32.6.
This research was conducted by Woohyeon Park, Jaeik Kim, and Sunghwan Cho of the ECE Department at SNU, with Professor Jaeyoung Do serving as the corresponding author. The study was also selected as an exemplary case of the government’s Supplementary Budget High-Performance Computing Support Program.
The research team plans to expand this work into next-generation multimodal medical foundation models that integrate medical imaging, clinical text, and patient data. Professor Do stated, “This research is meaningful because it goes beyond simply improving the performance of medical AI. It implements the actual clinical diagnostic process—detection, comparison, and explanation—inside the AI model itself. Through continued collaboration with hospitals and industry, we will work to develop trustworthy AI technologies that make tangible contributions to patient diagnosis and treatment.”
Source: https://ece.snu.ac.kr/ece/news?md=v&bbsidx=57783
Translated by: Changhoon Kang, English Editor of the Department of Electrical and Computer Engineering, changhoon27@snu.ac.kr
