collaborators

10 papers

cs.CV2026

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation

Dazhao Du, Shiyan Du, Jian Liu +8

Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation. Multimodal large language models (MLLMs…

cs.CV2026

AdaTok: Self-Budgeting Image Tokenization with Quality-Preserving Dynamic Tokens

Xiaocheng Lu, Yuxi Chen, Jie Zhang +5

Image tokenizers, from 2D grids to recent 1D sequences, typically encode every image with the same fixed number of tokens. Yet visual complexity is highly heterogeneous, so a unifo…

cs.CV2026

polyDAG: Polynomial Acyclicity Constraints for Efficient Continuous Causal Discovery in Visual Semantic Graphs

Wenhao Zhang, Ramin Ramezani, Tao Han +2

Modern image-analysis pipelines often convert images into structured semantic variables, such as facial attributes, object concepts, and scene descriptors. Learning directed depend…

cs.CV2026

Train the Agent, Not the Expert: Learning to Harness Heterogeneous Experts for Multi-Turn Visual Reasoning

Yaowu Fan, Tao Han, Dazhao Du +2

Recent progress in computer vision has produced a wide range of powerful specialized models for detection, segmentation, counting, and other visual tasks. However, these models are…

cs.CV2026

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning

Dazhao Du, Jian Liu, Jialong Qin +7

Video large language models (Video LLMs) achieve strong benchmark accuracy, yet often answer video questions through shortcuts such as single-frame cues and language priors rather…

cs.CV2026

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues

Dazhao Du, Liao Duan, Jian Liu +5

Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal large language models (MLLMs)…