5 papers
RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos?
Hongjie Zhou, Shiqin Wang, Haoyang Chen +5
Remote-sensing videos enable real-time observation of changes in target attributes, short-term activities, and scene evolution. They record motion, actions, interactions, and scene…
Position: The Systemic Lack of Agency in Visual Reasoning
Yizhao Huang, Haoyang Chen, Shiqin Wang +6
This paper argues that a systemic lack of Agency constrains the implicit reasoning capabilities of current Vision-Language Models (VLMs). Implicit reasoning refers to the ability t…
Any2Any: Unified Arbitrary Modality Translation for Remote Sensing
Haoyang Chen, Jing Zhang, Hebaixu Wang +7
Multi-modal remote sensing imagery provides complementary observations of the same geographic scene, yet such observations are frequently incomplete in practice. Existing cross-mod…
Heuristic Self-Paced Learning for Domain Adaptive Semantic Segmentation under Adverse Conditions
Shiqin Wang, Haoyang Chen, Huaizhou Huang +6
The learning order of semantic classes significantly impacts unsupervised domain adaptation for semantic segmentation, especially under adverse weather conditions. Most existing cu…
Subjective Camera 1.0: Bridging Human Cognition and Visual Reconstruction through Sequence-Aware Sketch-Guided Diffusion
Haoyang Chen, Dongfang Sun, Caoyuan Ma +4
We introduce the concept of a subjective camera to reconstruct meaningful moments that physical cameras fail to capture. We propose Subjective Camera 1.0, a framework for reconstru…