5 papers
RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos?
Hongjie Zhou, Shiqin Wang, Haoyang Chen +5
Remote-sensing videos enable real-time observation of changes in target attributes, short-term activities, and scene evolution. They record motion, actions, interactions, and scene…
Position: The Systemic Lack of Agency in Visual Reasoning
Yizhao Huang, Haoyang Chen, Shiqin Wang +6
This paper argues that a systemic lack of Agency constrains the implicit reasoning capabilities of current Vision-Language Models (VLMs). Implicit reasoning refers to the ability t…
Heuristic Self-Paced Learning for Domain Adaptive Semantic Segmentation under Adverse Conditions
Shiqin Wang, Haoyang Chen, Huaizhou Huang +6
The learning order of semantic classes significantly impacts unsupervised domain adaptation for semantic segmentation, especially under adverse weather conditions. Most existing cu…
Subjective Camera 1.0: Bridging Human Cognition and Visual Reconstruction through Sequence-Aware Sketch-Guided Diffusion
Haoyang Chen, Dongfang Sun, Caoyuan Ma +4
We introduce the concept of a subjective camera to reconstruct meaningful moments that physical cameras fail to capture. We propose Subjective Camera 1.0, a framework for reconstru…
Leader and Follower: Interactive Motion Generation under Trajectory Constraints
Runqi Wang, Caoyuan Ma, Jian Zhao +6
With the rapid advancement of game and film production, generating interactive motion from texts has garnered significant attention due to its potential to revolutionize content cr…