2 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.CV2023
No-frills Temporal Video Grounding: Multi-Scale Neighboring Attention and Zoom-in Boundary Detection
Qi Zhang, Sipeng Zheng, Qin Jin
Temporal video grounding (TVG) aims to retrieve the time interval of a language query from an untrimmed video. A significant challenge in TVG is the low "Semantic Noise Ratio (SNR)…
cs.CV2023★ 1 cited
Accommodating Audio Modality in CLIP for Multimodal Processing
Ludan Ruan, Anwen Hu, Yuqing Song +3
Multimodal processing has attracted much attention lately especially with the success of pre-training. However, the exploration has mainly focused on vision-language pre-training,…
cs.CV2022★ 2 cited
Exploring Anchor-based Detection for Ego4D Natural Language Query
Sipeng Zheng, Qi Zhang, Bei Liu +2
In this paper we provide the technique report of Ego4D natural language query challenge in CVPR 2022. Natural language query task is challenging due to the requirement of comprehen…