activity
20242026
most citedMotion Anything: Any to Motion Generation

1 citations · 1 across the 17 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2026

DOSE: Data Selection for Multi-Modal LLMs via Off-the-Shelf Models

Biao Wu, Yiwu Zhong, Meng Fang +1

High-quality and diverse multimodal data are essential for improving vision-language models (VLMs), yet existing datasets often contain noisy, redundant, and poorly aligned samples…

cs.CV2026

FlashSign: Pose-Free Guidance for Efficient Sign Language Video Generation

Liuzhou Zhang, Zeyu Zhang, Biao Wu +10

Sign language plays a crucial role in bridging communication gaps between the deaf and hard-of-hearing communities. However, existing sign language video generation models often re…

cs.CV2026

Transferring Physical Priors into Remote Sensing Segmentation via Large Language Models

Yuxi Lu, Kunqi Li, Zhidong Li +4

Semantic segmentation of remote sensing imagery is fundamental to Earth observation. Achieving accurate results requires integrating not only optical images but also physical varia…

cs.CV2025

UniVid: The Open-Source Unified Video Model

Jiabin Luo, Junhui Lin, Zeyu Zhang +4

Unified video modeling that combines generation and understanding capabilities is increasingly important but faces two key challenges: maintaining semantic faithfulness during flow…

cs.CV2025

StereoAdapter: Adapting Stereo Depth Estimation to Underwater Scenes

Zhengri Wu, Yiran Wang, Yu Wen +3

Underwater stereo depth estimation provides accurate 3D geometry for robotics tasks such as navigation, inspection, and mapping, offering metric depth from low-cost passive cameras…

cs.CV2025

VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery

Jinchao Ge, Tengfei Cheng, Biao Wu +7

Understanding cultural heritage artifacts such as ancient Greek pottery requires expert-level reasoning that remains challenging for current MLLMs due to limited domain-specific da…