1 citations · 1 across the 17 of their papers we have counts for
10 papers · 1 filter
DOSE: Data Selection for Multi-Modal LLMs via Off-the-Shelf Models
Biao Wu, Yiwu Zhong, Meng Fang +1
High-quality and diverse multimodal data are essential for improving vision-language models (VLMs), yet existing datasets often contain noisy, redundant, and poorly aligned samples…
FlashSign: Pose-Free Guidance for Efficient Sign Language Video Generation
Liuzhou Zhang, Zeyu Zhang, Biao Wu +10
Sign language plays a crucial role in bridging communication gaps between the deaf and hard-of-hearing communities. However, existing sign language video generation models often re…
Transferring Physical Priors into Remote Sensing Segmentation via Large Language Models
Yuxi Lu, Kunqi Li, Zhidong Li +4
Semantic segmentation of remote sensing imagery is fundamental to Earth observation. Achieving accurate results requires integrating not only optical images but also physical varia…
UniVid: The Open-Source Unified Video Model
Jiabin Luo, Junhui Lin, Zeyu Zhang +4
Unified video modeling that combines generation and understanding capabilities is increasingly important but faces two key challenges: maintaining semantic faithfulness during flow…
StereoAdapter: Adapting Stereo Depth Estimation to Underwater Scenes
Zhengri Wu, Yiran Wang, Yu Wen +3
Underwater stereo depth estimation provides accurate 3D geometry for robotics tasks such as navigation, inspection, and mapping, offering metric depth from low-cost passive cameras…
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
Jinchao Ge, Tengfei Cheng, Biao Wu +7
Understanding cultural heritage artifacts such as ancient Greek pottery requires expert-level reasoning that remains challenging for current MLLMs due to limited domain-specific da…