90 citations · 124 across the 14 of their papers we have counts for
6 papers · 1 filter
OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
Zhihong Chen, Xuehai Bai, Yang Shi +9
The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existi…
RaVL: Discovering and Mitigating Spurious Correlations in Fine-Tuned Vision-Language Models
Maya Varma, Jean-Benoit Delbrouck, Zhihong Chen +2
Fine-tuned vision-language models (VLMs) often capture spurious correlations between image features and textual attributes, resulting in degraded zero-shot performance at test time…
CheXalign: Preference fine-tuning in chest X-ray interpretation models without human feedback
Dennis Hein, Zhihong Chen, Sophie Ostmeier +8
Radiologists play a crucial role in translating medical images into actionable reports. However, the field faces staffing shortages and increasing workloads. While automated approa…
Merlin: A Computed Tomography Vision-Language Foundation Model and Dataset
Louis Blankemeier, Ashwin Kumar, Joseph Paul Cohen +37
The large volume of abdominal computed tomography (CT) scans coupled with the shortage of radiologists have intensified the need for automated medical image analysis tools. Previou…
A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation
Zhihong Chen, Maya Varma, Justin Xu +20
Over 1.4 billion chest X-rays (CXRs) are performed annually due to their cost-effectiveness as an initial diagnostic test. This scale of radiological studies provides a significant…
Exploiting Low-confidence Pseudo-labels for Source-free Object Detection
Zhihong Chen, Zilei Wang, Yixin Zhang
Source-free object detection (SFOD) aims to adapt a source-trained detector to an unlabeled target domain without access to the labeled source data. Current SFOD methods utilize a…