5 papers
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching
Mengxiao Tian, Xinxiao Wu, Shuo Yang
Driven by large-scale contrastive vision-language pre-trained models such as CLIP, recent advancements in the image-text matching task have achieved remarkable success in represent…
LLM-powered Query Expansion for Enhancing Boundary Prediction in Language-driven Action Localization
Zirui Shang, Xinxiao Wu, Shuo Yang
Language-driven action localization in videos requires not only semantic alignment between language query and video segment, but also prediction of action boundaries. However, the…
METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection
Yongqi Wang, Xinxiao Wu, Shuo Yang
Open-vocabulary video visual relationship detection aims to detect objects and their relationships in videos without being restricted by predefined object or relationship categorie…
End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting
Yongqi Wang, Xinxiao Wu, Shuo Yang +1
Open-vocabulary video visual relationship detection aims to expand video visual relationship detection beyond annotated categories by detecting unseen relationships between both se…
Video Summarization using Denoising Diffusion Probabilistic Model
Zirui Shang, Yubo Zhu, Hongxi Li +2
Video summarization aims to eliminate visual redundancy while retaining key parts of video to construct concise and comprehensive synopses. Most existing methods use discriminative…