collaborators

5 papers

cs.CV2025

GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents

Xianhang Ye, Yiqing Li, Wei Dai +8

Existing GUI grounding methods often struggle with fine-grained localization in high-resolution screenshots. To address this, we propose GUI-ARP, a novel framework that enables ada…

cs.AI2025

ViDove: A Translation Agent System with Multimodal Context and Memory-Augmented Reasoning

Yichen Lu, Wei Dai, Jiaen Liu +9

LLM-based translation agents have achieved highly human-like translation results and are capable of handling longer and more complex contexts with greater efficiency. However, they…

cs.HC2025

A Human-Centered Approach to Identifying Promises, Risks, & Challenges of Text-to-Image Generative AI in Radiology

Katelyn Morrison, Arpit Mathur, Aidan Bradshaw +7

As text-to-image generative models rapidly improve, AI researchers are making significant advances in developing domain-specific models capable of generating complex medical imager…

cs.LG2025

QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training

Wei Dai, Peilin Chen, Chanakya Ekbote +1

Clinical decision-making routinely demands reasoning over heterogeneous data, yet existing multimodal language models (MLLMs) remain largely vision-centric and fail to generalize a…

cs.LG2025

CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models

Wei Dai, Peilin Chen, Malinda Lu +4

Recent advances in clinical AI have enabled remarkable progress across many clinical domains. However, existing benchmarks and models are primarily limited to a small set of modali…