10 papers
Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation
Dazhao Du, Shiyan Du, Jian Liu +8
Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation. Multimodal large language models (MLLMs…
AdaTok: Self-Budgeting Image Tokenization with Quality-Preserving Dynamic Tokens
Xiaocheng Lu, Yuxi Chen, Jie Zhang +5
Image tokenizers, from 2D grids to recent 1D sequences, typically encode every image with the same fixed number of tokens. Yet visual complexity is highly heterogeneous, so a unifo…
polyDAG: Polynomial Acyclicity Constraints for Efficient Continuous Causal Discovery in Visual Semantic Graphs
Wenhao Zhang, Ramin Ramezani, Tao Han +2
Modern image-analysis pipelines often convert images into structured semantic variables, such as facial attributes, object concepts, and scene descriptors. Learning directed depend…
Train the Agent, Not the Expert: Learning to Harness Heterogeneous Experts for Multi-Turn Visual Reasoning
Yaowu Fan, Tao Han, Dazhao Du +2
Recent progress in computer vision has produced a wide range of powerful specialized models for detection, segmentation, counting, and other visual tasks. However, these models are…
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning
Dazhao Du, Jian Liu, Jialong Qin +7
Video large language models (Video LLMs) achieve strong benchmark accuracy, yet often answer video questions through shortcuts such as single-frame cues and language priors rather…
MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues
Dazhao Du, Liao Duan, Jian Liu +5
Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal large language models (MLLMs)…