10 papers
Recognition-Conditioned Reasoning: A Training-Free Multimodal-LLM Pipeline for Fine-Grained Micro-Action Understanding
Fengshun Wang, Jin'ang Han, Zhigang Tu
Micro-actions are subtle, short, low-amplitude body movements, such as a fidgeting hand or a slight head tilt, that humans perform with little conscious intent yet that reliably le…
DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding
Zhengbo Zhang, Mark He Huang, Zhigang Tu +1
Zero-shot video temporal grounding (VTG) localizes events in untrimmed videos from natural language queries without task-specific training. Existing methods rely on frame-query fea…
Frequency-Enhanced Diffusion Models: Curriculum-Guided Semantic Alignment for Zero-Shot Skeleton Action Recognition
Yuxi Zhou, Zhengbo Zhang, Jingyu Pan +2
Human action recognition is pivotal in computer vision, with applications ranging from surveillance to human-robot interaction. Despite the effectiveness of supervised skeleton-bas…
Masked Diffusion Vision-Language Models for Temporal Action Localization
Fengshun Wang, Zhengbo Zhang, Zhigang Tu
Temporal action localization (TAL) requires recognizing the target event and localizing its start and end times precisely in untrimmed videos. Recent vision-language formulations i…
Leveraging Text-to-Image Diffusion Models for Unsupervised Visual Object Tracking
Zhengbo Zhang, Zhigang Tu, Junsong Yuan +2
Unsupervised visual object tracking is a challenging task that requires following arbitrary targets in videos without training on ground-truth annotations. Despite considerable pro…
UAV-OVO: Out-of-Viewpoint Generalization in UAV Action Recognition
Yu Xia, Zhengbo Zhang, Shuaihu Zhang +1
UAV action recognition faces a deployment shift that standard benchmarks often obscure: a model trained on UAV footage captured from low-depression viewpoints may be required to re…