papers

Publications (25)

cs.HC2026

Not Another EHR: Reimagining Physician Information Needs with Generative AI Technology

Ruican Zhong, Jiachen Li, Gary Hsieh +15

Electronic health records (EHRs) have improved data accessibility but have also introduced cognitive burden for physicians, given the sheer volume and complexity of the data involv…

cs.CV2019

Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC)

Noel Codella, Veronica Rotemberg, Philipp Tschandl +9

This work summarizes the results of the largest skin image analysis challenge in the world, hosted by the International Skin Imaging Collaboration (ISIC), a global partnership that…

cs.LG2020

P2L: Predicting Transfer Learning for Images and Semantic Relations

Bishwaranjan Bhattacharjee, John R. Kender, Matthew Hill +7

Transfer learning enhances learning across tasks, by leveraging previously learned representations -- if they are properly chosen. We describe an efficient method to accurately est…

cs.CV2016

Deep Learning Ensembles for Melanoma Recognition in Dermoscopy Images

Noel Codella, Quoc-Bao Nguyen, Sharath Pankanti +4

Melanoma is the deadliest form of skin cancer. While curable with early detection, only highly trained specialists are capable of accurately recognizing the disease. As expertise i…

cs.CV2021

CvT: Introducing Convolutions to Vision Transformers

Haiping Wu, Bin Xiao, Noel Codella +4

We present in this paper a new architecture, named Convolutional vision Transformer (CvT), that improves Vision Transformer (ViT) in performance and efficiency by introducing convo…

cs.CL2025

Closing the Performance Gap Between AI and Radiologists in Chest X-Ray Reporting

Harshita Sharma, Maxwell C. Reynolds, Valentina Salvatelli +26

AI-assisted report generation offers the opportunity to reduce radiologists' workload stemming from expanded screening guidelines, complex cases and workforce shortages, while main…

cs.CV2026

UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding

Rui Sun, Zhecan Wang, Haoxuan You +3

Vision-language tasks, such as VQA, SNLI-VE, and VCR are challenging because they require the model's reasoning ability to understand the semantics of the visual world and natural…

cs.LG2025

Demo: Healthcare Agent Orchestrator (HAO) for Patient Summarization in Molecular Tumor Boards

Matthias Blondeel, Noel Codella, Sam Preston +9

Molecular Tumor Boards (MTBs) are multidisciplinary forums where oncology specialists collaboratively assess complex patient cases to determine optimal treatment strategies. A cent…

cs.CV2022

Multimodal Adaptive Distillation for Leveraging Unimodal Encoders for Vision-Language Tasks

Zhecan Wang, Noel Codella, Yen-Chun Chen +8

Cross-modal encoders for vision-language (VL) tasks are often pretrained with carefully curated vision-language datasets. While these datasets reach an order of 10 million samples,…

eess.IV2020

A Patient-Centric Dataset of Images and Metadata for Identifying Melanomas Using Clinical Context

Veronica Rotemberg, Nicholas Kurtansky, Brigid Betz-Stablein +21

Prior skin image datasets have not addressed patient-level information obtained from multiple skin lesions from the same patient. Though artificial intelligence classification algo…

cs.CV2024

Fully Authentic Visual Question Answering Dataset from Online Communities

Chongyan Chen, Mengchen Liu, Noel Codella +3

Visual Question Answering (VQA) entails answering questions about images. We introduce the first VQA dataset in which all contents originate from an authentic use case. Sourced fro…

cs.CV2023

Streaming Video Model

Yucheng Zhao, Chong Luo, Chuanxin Tang +3

Video understanding tasks have traditionally been modeled by two separate architectures, specially tailored for two distinct tasks. Sequence-based video tasks, such as action recog…

cs.CV2021

RegionCLIP: Region-based Language-Image Pretraining

Yiwu Zhong, Jianwei Yang, Pengchuan Zhang +8

Contrastive language-image pretraining (CLIP) using image-text pairs has achieved impressive results on image classification in both zero-shot and transfer learning settings. Howev…

cs.CV2022

Learning Visual Representation from Modality-Shared Contrastive Language-Image Pre-training

Haoxuan You, Luowei Zhou, Bin Xiao +5

Large-scale multi-modal contrastive pre-training has demonstrated great utility to learn transferable features for a range of downstream tasks by mapping multiple modalities into a…

cs.CV2021

Florence: A New Foundation Model for Computer Vision

Lu Yuan, Dongdong Chen, Yi-Ling Chen +20

Automated visual understanding of our diverse and open world demands computer vision models to generalize well with minimal customization for specific tasks, similar to human visio…

cs.CV2025

Exploring scalable medical image encoders beyond text supervision

Fernando Pérez-García, Harshita Sharma, Sam Bond-Taylor +12

Language-supervised pre-training has proven to be a valuable method for extracting semantically meaningful features from images, serving as a foundational element in multimodal sys…

eess.IV2024

Generative Enhancement for 3D Medical Images

Lingting Zhu, Noel Codella, Dongdong Chen +3

The limited availability of 3D medical image datasets, due to privacy concerns and high collection or annotation costs, poses significant challenges in the field of medical imaging…

cs.CV2022

DaViT: Dual Attention Vision Transformers

Mingyu Ding, Bin Xiao, Noel Codella +3

In this work, we introduce Dual Attention Vision Transformers (DaViT), a simple yet effective vision transformer architecture that is able to capture global context while maintaini…

cs.CL2023

i-Code V2: An Autoregressive Generation Framework over Vision, Language, and Speech Data

Ziyi Yang, Mahmoud Khademi, Yichong Xu +16

The convergence of text, visual, and audio data is a key step towards human-like artificial intelligence, however the current Vision-Language-Speech landscape is dominated by encod…

cs.CL2024

MAIRA-1: A specialised large multimodal model for radiology report generation

Stephanie L. Hyland, Shruthi Bannur, Kenza Bouzid +12

We present a radiology-specific multimodal model for the task for generating radiological reports from chest X-rays (CXRs). Our work builds on the idea that large language model(s)…

cs.CV2022

CLIP-TD: CLIP Targeted Distillation for Vision-Language Tasks

Zhecan Wang, Noel Codella, Yen-Chun Chen +7

Contrastive language-image pretraining (CLIP) links vision and language modalities into a unified embedding space, yielding the tremendous potential for vision-language (VL) tasks.…

cs.CV2023

Dataset Bias Mitigation in Multiple-Choice Visual Question Answering and Beyond

Zhecan Wang, Long Chen, Haoxuan You +6

Vision-language (VL) understanding tasks evaluate models' comprehension of complex visual scenes through multiple-choice questions. However, we have identified two dataset biases t…

cs.LG2022

i-Code: An Integrative and Composable Multimodal Learning Framework

Ziyi Yang, Yuwei Fang, Chenguang Zhu +17

Human intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to…

cs.CV2025

DermaVQA-DAS: Dermatology Assessment Schema (DAS) & Datasets for Closed-Ended Question Answering & Segmentation in Patient-Generated Dermatology Images

Wen-wai Yim, Yujuan Fu, Asma Ben Abacha +4

Recent advances in dermatological image analysis have been driven by large-scale annotated datasets; however, most existing benchmarks focus on dermatoscopic images and lack patien…

cs.LG2016

Learning Representations from EEG with Deep Recurrent-Convolutional Neural Networks

Pouya Bashivan, Irina Rish, Mohammed Yeasin +1

One of the challenges in modeling cognitive events from electroencephalogram (EEG) data is finding representations that are invariant to inter- and intra-subject differences, as we…