8 papers
CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming
Ruixun Liu, Lingyu Zhang, Lanxuan Xue +3
Humans can effortlessly reason about scenes across different viewpoints, yet it remains unclear whether Vision-Language Models (VLMs) possess similar cross-view spatial abilities.…
MEval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks
Jie Huang, Ruixun Liu, Sirui Sun +4
As multi-modal models advance towards long-form video understanding, memory emerges as a critical capability. Despite substantial efforts in developing video datasets and benchmark…
ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks
Ruixun Liu, Bowen Fu, Jiayi Song +7
Ultra-high-resolution (UHR) remote sensing (RS) images offer rich fine-grained information but also present challenges in effective processing. Existing dynamic resolution and toke…
Annotation-Free Open-Vocabulary Segmentation for Remote-Sensing Images
Kaiyu Li, Xiangyong Cao, Ruixun Liu +4
Semantic segmentation of remote sensing (RS) images is pivotal for comprehensive Earth observation, but the demand for interpreting new object categories, coupled with the high exp…
Flow Diverse and Efficient: Learning Momentum Flow Matching via Stochastic Velocity Field Sampling
Zhiyuan Ma, Ruixun Liu, Sixian Liu +2
Recently, the rectified flow (RF) has emerged as the new state-of-the-art among flow-based diffusion models due to its high efficiency advantage in straight path sampling, especial…
SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing Images
Kaiyu Li, Ruixun Liu, Xiangyong Cao +4
Remote sensing image plays an irreplaceable role in fields such as agriculture, water resources, military, and disaster relief. Pixel-level interpretation is a critical aspect of r…