6 papers
4DPChat: Towards Dynamic Point Cloud Understanding with Failure-Aware Bootstrapping
Xindan Zhang, Weilong Yan, Yufei Shi +5
Point clouds provide a compact and expressive representation of 3D objects, and have recently been integrated into multimodal large language models (MLLMs). However, existing metho…
FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO
Fufangchen Zhao, Songbai Tan, Xuerui Qiu +7
Existing video large language models (VLLMs) primarily leverage prompt agnostic visual encoders, which extract untargeted facial representations without awareness of the queried in…
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration
Kerui Chen, Jinglu Wang, Jianrong Zhang +3
Multi-modal large language models (MLLMs) advance vision language understanding but face inherent limitations in long-video tasks due to bounded perception context budgets. Existin…
RS-WorldModel: a Unified Model for Remote Sensing Understanding and Future Sense Forecasting
Linrui Xu, Zhongan Wang, Fei Shen +4
Remote sensing world models aim to both explain observed changes and forecast plausible futures, two tasks that share spatiotemporal priors. Existing methods, however, typically ad…
UniFace: A Unified Fine-grained Face Understanding and Generation Model
Junzhe Li, Sifan Zhou, Liya Guo +9
Unified multimodal models (UMMs) have emerged as a powerful paradigm in fundamental cross-modality research, demonstrating significant potential in both image understanding and gen…
SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System
Zhiyu Xu, Weilong Yan, Yufei Shi +5
Recent advancements in multimodal large language models (MLLMs) and video agent systems have significantly improved general video understanding. However, when applied to scientific…