1 citations · 1 across the 10 of their papers we have counts for
11 papers
DC-Motion: Decoupling Structure and Details via Discrete-Continuous Tokens for Human Motion Generation
Hequan Wang, Xuean Chen, Jiaxu Zhang +2
Text-to-motion generation requires modeling both global action structure and fine-grained motion dynamics from natural language. Existing approaches typically rely on either contin…
DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding
Zhengbo Zhang, Mark He Huang, Zhigang Tu +1
Zero-shot video temporal grounding (VTG) localizes events in untrimmed videos from natural language queries without task-specific training. Existing methods rely on frame-query fea…
Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning
Zhengbo Zhang, Changtao Miao, Jinbo Su +10
Multimodal large language models (MLLMs) have demonstrated impressive capabilities in many visual tasks, but they often struggle with factual grounding when confronted with complex…
Frequency-Enhanced Diffusion Models: Curriculum-Guided Semantic Alignment for Zero-Shot Skeleton Action Recognition
Yuxi Zhou, Zhengbo Zhang, Jingyu Pan +2
Human action recognition is pivotal in computer vision, with applications ranging from surveillance to human-robot interaction. Despite the effectiveness of supervised skeleton-bas…
Masked Diffusion Vision-Language Models for Temporal Action Localization
Fengshun Wang, Zhengbo Zhang, Zhigang Tu
Temporal action localization (TAL) requires recognizing the target event and localizing its start and end times precisely in untrimmed videos. Recent vision-language formulations i…
Leveraging Text-to-Image Diffusion Models for Unsupervised Visual Object Tracking
Zhengbo Zhang, Zhigang Tu, Junsong Yuan +2
Unsupervised visual object tracking is a challenging task that requires following arbitrary targets in videos without training on ground-truth annotations. Despite considerable pro…