6 papers
Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation
Dazhao Du, Shiyan Du, Jian Liu +8
Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation. Multimodal large language models (MLLMs…
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning
Dazhao Du, Jian Liu, Jialong Qin +7
Video large language models (Video LLMs) achieve strong benchmark accuracy, yet often answer video questions through shortcuts such as single-frame cues and language priors rather…
MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues
Dazhao Du, Liao Duan, Jian Liu +5
Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal large language models (MLLMs)…
PPAI: Enabling Personalized LLM Agent Interoperability for Collaborative Edge Intelligence
Zile Wang, Qianli Liu, Kaibin Guo +4
Deploying large language model (LLM) on edge device enables personalized LLM agents for various users. The growing availability of diverse personalized agents presents a unique opp…
What You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic Curiosity
Haoxi Li, Qinglin Hou, Jianfei Ma +6
To navigate partially observable visual environments, recent VLM agents increasingly internalize world modeling capabilities into their policies via explicit CoT reasoning, enablin…
Predicting the Future by Retrieving the Past
Dazhao Du, Tao Han, Song Guo
Deep learning models such as MLP, Transformer, and TCN have achieved remarkable success in univariate time series forecasting, typically relying on sliding window samples from hist…