4 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CV2025
LLM-powered Query Expansion for Enhancing Boundary Prediction in Language-driven Action Localization
Zirui Shang, Xinxiao Wu, Shuo Yang
Language-driven action localization in videos requires not only semantic alignment between language query and video segment, but also prediction of action boundaries. However, the…
cs.CV2024
Video Summarization using Denoising Diffusion Probabilistic Model
Zirui Shang, Yubo Zhu, Hongxi Li +2
Video summarization aims to eliminate visual redundancy while retaining key parts of video to construct concise and comprehensive synopses. Most existing methods use discriminative…
cs.CV2024★ 4 cited
Data-free Multi-label Image Recognition via LLM-powered Prompt Tuning
Shuo Yang, Zirui Shang, Yongqi Wang +4
This paper proposes a novel framework for multi-label image recognition without any training data, called data-free framework, which uses knowledge of pre-trained Large Language Mo…