activity
20242026
collaborators

6 papers

cs.AI2026

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning

Yongxin Wang, Ruizhe Zhou, Yueling Tang +4

Multimodal large language models increasingly reason over screenshots and documents where the task itself may be written in pixels. Yet benchmarks usually place questions in text,…

cs.RO2025

GLaD: Geometric Latent Distillation for Vision-Language-Action Models

Minghao Guo, Meng Cao, Jiachen Tao +5

Most existing Vision-Language-Action (VLA) models rely primarily on RGB information, while ignoring geometric cues crucial for spatial reasoning and manipulation. In this work, we…

cs.CV2025

Medical Report Generation Is A Multi-label Classification Problem

Yijian Fan, Zhenbang Yang, Rui Liu +2

Medical report generation is a critical task in healthcare that involves the automatic creation of detailed and accurate descriptions from medical images. Traditionally, this task…

cs.CV2025

NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning

Bingqian Lin, Yunshuang Nie, Ziming Wei +6

Vision-and-Language Navigation (VLN), as a crucial research problem of Embodied AI, requires an embodied agent to navigate through complex 3D environments following natural languag…

cs.CV2025

Sitcom-Crafter: A Plot-Driven Human Motion Generation System in 3D Scenes

Jianqi Chen, Panwen Hu, Xiaojun Chang +3

Recent advancements in human motion synthesis have focused on specific types of motions, such as human-scene interaction, locomotion or human-human interaction, however, there is a…

cs.CV2024

HC-LLM: Historical-Constrained Large Language Models for Radiology Report Generation

Tengfei Liu, Jiapu Wang, Yongli Hu +5

Radiology report generation (RRG) models typically focus on individual exams, often overlooking the integration of historical visual or textual data, which is crucial for patient f…