2 citations · 3 across the 6 of their papers we have counts for
4 papers · 1 filter
Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?
Xiwei Liu, Yulong Li, Xinlin Zhuang +5
Vision-language models (VLMs) can answer spatial questions, yet the mechanisms connecting object grounding to spatial reasoning remain poorly understood. It is underexplored whethe…
Attention-Guided Flow-Matching for Sparse 3D Geological Generation
Zhixiang Lu, Mengqi Han, Peixin Guo +4
Constructing high-resolution 3D geological models from sparse 1D borehole and 2D surface data is a highly ill-posed inverse problem. Traditional heuristic and implicit modeling met…
Semantic-Topological Graph Reasoning for Language-Guided Pulmonary Screening
Chenyu Xue, Yiran Liu, Mian Zhou +2
Medical image segmentation driven by free-text clinical instructions is a critical frontier in computer-aided diagnosis. However, existing multimodal and foundation models struggle…
Beyond Words: AuralLLM and SignMST-C for Sign Language Production and Bidirectional Accessibility
Yulong Li, Yuxuan Zhang, Feilong Tang +10
Sign language is the primary communication mode for 72 million hearing-impaired individuals worldwide, necessitating effective bidirectional Sign Language Production and Sign Langu…