activity
20242026
collaborators

15 papers

cs.CV2026

Learning to Use Imagination: Progress-Conditioned Future Utilization for World Action Models

Yijie Zhu, Zitong Yu, Wei Li +4

World Action Models (WAMs) extend Vision-Language-Action (VLA) models by incorporating future visual dynamics into action generation. However, existing WAMs often utilize imagined…

cs.CV2026

PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models

Zihan Song, Shuo Ye, Bo Zhao +4

Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-duration frames remains a hindranc…

cs.RO2026

AC-VLA: Robust Out-of-Distribution Action Execution via Compositional Learning

Xiaojiang Peng, Kai Peng, Jie Lu +3

Vision-Language-Action (VLA) models excel at end-to-end robotic manipulation but struggle with out-of-distribution (OOD) generalization when familiar sub-tasks are recombined in un…

cs.CV2026

GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs

Taorui Wang, Wei Xia, Hui Ma +5

Micro-gesture recognition demands the detection of fleeting, spatially localized movements that are frequently overwhelmed by dominant static appearances and background noise. Whil…

cs.CV2026

RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning

Yelin Wang, Zijia Song, Shuo Ye +6

Remote Sensing Image Change Captioning (RSICC) aims to describe changes between bi-temporal remote sensing images and holds significant research and application value. However, mos…

cs.CV2026

OracleAnalyser: Analysing Implicit Semantics of Oracle Bone Scripts through MLLMs with Post-training

Zijia Song, Yelin Wang, Zhengyi Ma +5

With the advancement of artificial intelligence, research on oracle bone scripts has entered a new era. However, existing methods and benchmarks remain largely confined to recognit…