2 papers
cs.CV2024
Convex Combination Consistency between Neighbors for Weakly-supervised Action Localization
Qinying Liu, Zilei Wang, Ruoxi Chen +1
Weakly-supervised temporal action localization (WTAL) intends to detect action instances with only weak supervision, e.g., video-level labels. The current~\textit{de facto} pipelin…
cs.CV2024
TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification
Qinying Liu, Wei Wu, Kecheng Zheng +6
The crux of learning vision-language models is to extract semantically aligned information from visual and linguistic data. Existing attempts usually face the problem of coarse ali…