3 papers
cs.CV2026
Video Understanding: From Geometry and Semantics to Unified Models
Zhaochong An, Zirui Li, Mingqiao Ye +9
Video understanding aims to enable models to perceive, reason about, and interact with the dynamic visual world. In contrast to image understanding, video understanding inherently…
cs.AI2025
ChatMotion: A Multimodal Multi-Agent for Human Motion Analysis
Lei Li, Sen Jia, Jianhao Wang +4
Advancements in Multimodal Large Language Models (MLLMs) have improved human motion understanding. However, these models remain constrained by their "instruct-only" nature, lacking…
cs.AI2024
Adaptive Masking Enhances Visual Grounding
Sen Jia, Lei Li
In recent years, zero-shot and few-shot learning in visual grounding have garnered considerable attention, largely due to the success of large-scale vision-language pre-training on…