Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
R^3-VQA: "Read the Room" by Video Social Reasoning
Lixing Niu, Jiapeng Li, Xingping Yu +6
"Read the room" is a significant social reasoning capability in human daily life. Humans can infer others' mental states from subtle social cues. Previous social reasoning tasks an…
cs.CV2025
DMPT: Decoupled Modality-aware Prompt Tuning for Multi-modal Object Re-identification
Minghui Lin, Shu Wang, Xiang Wang +4
Current multi-modal object re-identification approaches based on large-scale pre-trained backbones (i.e., ViT) have displayed remarkable progress and achieved excellent performance…
cs.CV2025
Exploring the Evolution of Physics Cognition in Video Generation: A Survey
Minghui Lin, Xiang Wang, Yishan Wang +8
Recent advancements in video generation have witnessed significant progress, especially with the rapid advancement of diffusion models. Despite this, their deficiencies in physical…