activity
20242026
collaborators

8 papers

cs.CV2026

CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming

Ruixun Liu, Lingyu Zhang, Lanxuan Xue +3

Humans can effortlessly reason about scenes across different viewpoints, yet it remains unclear whether Vision-Language Models (VLMs) possess similar cross-view spatial abilities.…

cs.CV2026

MEval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks

Jie Huang, Ruixun Liu, Sirui Sun +4

As multi-modal models advance towards long-form video understanding, memory emerges as a critical capability. Despite substantial efforts in developing video datasets and benchmark…

cs.CV2025

ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks

Ruixun Liu, Bowen Fu, Jiayi Song +7

Ultra-high-resolution (UHR) remote sensing (RS) images offer rich fine-grained information but also present challenges in effective processing. Existing dynamic resolution and toke…

cs.CV2025

Annotation-Free Open-Vocabulary Segmentation for Remote-Sensing Images

Kaiyu Li, Xiangyong Cao, Ruixun Liu +4

Semantic segmentation of remote sensing (RS) images is pivotal for comprehensive Earth observation, but the demand for interpreting new object categories, coupled with the high exp…

cs.CV2025

Flow Diverse and Efficient: Learning Momentum Flow Matching via Stochastic Velocity Field Sampling

Zhiyuan Ma, Ruixun Liu, Sixian Liu +2

Recently, the rectified flow (RF) has emerged as the new state-of-the-art among flow-based diffusion models due to its high efficiency advantage in straight path sampling, especial…

cs.CV2024

SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing Images

Kaiyu Li, Ruixun Liu, Xiangyong Cao +4

Remote sensing image plays an irreplaceable role in fields such as agriculture, water resources, military, and disaster relief. Pixel-level interpretation is a critical aspect of r…