8 papers
LaVIDE: Language-Prompted Satellite Change Detection via Map-Image Alignment
Shuguo Jiang, Fang Xu, Chuandong Liu +6
Remote sensing change detection based on a map reference and an up-to-date image boosts timely observation of the Earth's surface when earlier images are lacking for comparison. Ho…
RAM: Recover Any 3D Human Motion in-the-Wild
Sen Jia, Ning Zhu, Jinqin Zhong +4
RAM incorporates a motion-aware semantic tracker with adaptive Kalman filtering to achieve robust identity association under severe occlusions and dynamic interactions. A memory-au…
Learning an Efficient Optimizer via Hybrid-Policy Sub-Trajectory Balance
Yunchuan Guan, Yu Liu, Ke Zhou +8
Recent advances in generative modeling enable neural networks to generate weights without relying on gradient-based optimization. However, current methods are limited by issues of…
COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework
Xin Dong, Sen Jia, Ming Rui Wang +4
Recently, with the emergence of recent Multimodal Large Language Model (MLLM) technology, it has become possible to exploit its video understanding capability on different classifi…
Human Motion Instruction Tuning
Lei Li, Sen Jia, Jianhao Wang +6
This paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning ap…
ChatMotion: A Multimodal Multi-Agent for Human Motion Analysis
Lei Li, Sen Jia, Jianhao Wang +4
Advancements in Multimodal Large Language Models (MLLMs) have improved human motion understanding. However, these models remain constrained by their "instruct-only" nature, lacking…