1 citations · 1 across the 3 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
CheXTemporal: A Dataset for Temporally-Grounded Reasoning in Chest Radiography
Eva Prakash, Yunhe Gao, Chong Wang +10
Chest radiograph interpretation requires temporal reasoning over prior and current studies, yet most vision-language models are trained on static image-report pairs and lack explic…
cs.CV2026
LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?
Kechen Fang, Yihua Qin, Chongyi Wang +3
Visual encoding constitutes a major computational bottleneck in Multimodal Large Language Models (MLLMs), especially for high-resolution image inputs. The prevailing practice typic…
cs.CV2024
MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Yuan Yao, Tianyu Yu, Ao Zhang +20
The recent surge of Multimodal Large Language Models (MLLMs) has fundamentally reshaped the landscape of AI research and industry, shedding light on a promising path toward the nex…