collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation

Dazhao Du, Shiyan Du, Jian Liu +8

Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation. Multimodal large language models (MLLMs…

cs.CV2026

Train the Agent, Not the Expert: Learning to Harness Heterogeneous Experts for Multi-Turn Visual Reasoning

Yaowu Fan, Tao Han, Dazhao Du +2

Recent progress in computer vision has produced a wide range of powerful specialized models for detection, segmentation, counting, and other visual tasks. However, these models are…

cs.CV2026

WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models

Bohai Gu, Taiyi Wu, Yueyang Yuan +9

Recent video-based world models have made pixel-space environments interactive at the camera level: users can navigate viewpoints while the model generates coherent visual continua…

cs.CV2026

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning

Dazhao Du, Jian Liu, Jialong Qin +7

Video large language models (Video LLMs) achieve strong benchmark accuracy, yet often answer video questions through shortcuts such as single-frame cues and language priors rather…

cs.CV2026

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues

Dazhao Du, Liao Duan, Jian Liu +5

Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal large language models (MLLMs)…

cs.CV2026

Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion

Bohai Gu, Taiyi Wu, Dazhao Du +5

Video object insertion is fundamental to video editing, yet existing diffusion methods often produce visually plausible but physically inconsistent results. We present Place-it-R1,…