6 papers
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
Hangjie Yuan, Yichen Qian, Zhiwei Tang +21
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric chall…
Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts
Hongkun Pan, Yuwei Wu, Wanyi Hong +8
Multimodal large language models (MLLMs) have shown considerable potential in chart understanding and reasoning tasks. However, they still struggle with high information density (H…
V-Zero: Self-Improving Multimodal Reasoning with Zero Annotation
Han Wang, Yi Yang, Jingyuan Hu +2
Recent advances in multimodal learning have significantly enhanced the reasoning capabilities of vision-language models (VLMs). However, state-of-the-art approaches rely heavily on…
R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Yi Yang, Xiaoxuan He, Hongkun Pan +9
Large Language Models have demonstrated remarkable reasoning capability in complex textual tasks. However, multimodal reasoning, which requires integrating visual and textual infor…
Exploring the Interplay Between Video Generation and World Models in Autonomous Driving: A Survey
Ao Fu, Yi Zhou, Tao Zhou +5
World models and video generation are pivotal technologies in the domain of autonomous driving, each playing a critical role in enhancing the robustness and reliability of autonomo…
SPROUT: an Interactive Authoring Tool for Generating Programming Tutorials with the Visualization of Large Language Models
Yihan Liu, Zhen Wen, Luoxuan Weng +3
The rapid development of large language models (LLMs), such as ChatGPT, has revolutionized the efficiency of creating programming tutorials. LLMs can be instructed with text prompt…