6 papers
When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning
Yongxin Wang, Ruizhe Zhou, Yueling Tang +4
Multimodal large language models increasingly reason over screenshots and documents where the task itself may be written in pixels. Yet benchmarks usually place questions in text,…
GLaD: Geometric Latent Distillation for Vision-Language-Action Models
Minghao Guo, Meng Cao, Jiachen Tao +5
Most existing Vision-Language-Action (VLA) models rely primarily on RGB information, while ignoring geometric cues crucial for spatial reasoning and manipulation. In this work, we…
Medical Report Generation Is A Multi-label Classification Problem
Yijian Fan, Zhenbang Yang, Rui Liu +2
Medical report generation is a critical task in healthcare that involves the automatic creation of detailed and accurate descriptions from medical images. Traditionally, this task…
NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning
Bingqian Lin, Yunshuang Nie, Ziming Wei +6
Vision-and-Language Navigation (VLN), as a crucial research problem of Embodied AI, requires an embodied agent to navigate through complex 3D environments following natural languag…
Sitcom-Crafter: A Plot-Driven Human Motion Generation System in 3D Scenes
Jianqi Chen, Panwen Hu, Xiaojun Chang +3
Recent advancements in human motion synthesis have focused on specific types of motions, such as human-scene interaction, locomotion or human-human interaction, however, there is a…
HC-LLM: Historical-Constrained Large Language Models for Radiology Report Generation
Tengfei Liu, Jiapu Wang, Yongli Hu +5
Radiology report generation (RRG) models typically focus on individual exams, often overlooking the integration of historical visual or textual data, which is crucial for patient f…