2 papers
cs.CV2026
Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark
Seng Nam Chen, Hao Chen, Chenglam Ho +4
Long video understanding (LVU) remains a core challenge in multimodal learning. Although recent vision-language models (VLMs) have made notable progress, existing benchmarks mainly…
cs.CV2025
FusDreamer: Label-efficient Remote Sensing World Model for Multimodal Data Classification
Jinping Wang, Weiwei Song, Hao Chen +2
World models significantly enhance hierarchical understanding, improving data integration and learning efficiency. To explore the potential of the world model in the remote sensing…