93 citations · 106 across the 6 of their papers we have counts for
9 papers
Towards Debiasing Temporal Sentence Grounding in Video
Hao Zhang, Aixin Sun, Wei Jing +1
The temporal sentence grounding in video (TSGV) task is to locate a temporal moment from an untrimmed video, to match a language query, i.e., a sentence. Without considering bias i…
Domain Generalization for Vision-based Driving Trajectory Generation
Yunkai Wang, Dongkun Zhang, Yuxiang Cui +5
One of the challenges in vision-based driving trajectory generation is dealing with out-of-distribution scenarios. In this paper, we propose a domain generalization method for visi…
Video Corpus Moment Retrieval with Contrastive Learning
Hao Zhang, Aixin Sun, Wei Jing +4
Given a collection of untrimmed and unsegmented videos, video corpus moment retrieval (VCMR) is to retrieve a temporal moment (i.e., a fraction of a video) that semantically corres…
Natural Language Video Localization: A Revisit in Span-based Question Answering Framework
Hao Zhang, Aixin Sun, Wei Jing +3
Natural Language Video Localization (NLVL) aims to locate a target moment from an untrimmed video that semantically corresponds to a text query. Existing approaches mainly solve th…
Context Modeling with Evidence Filter for Multiple Choice Question Answering
Sicheng Yu, Hao Zhang, Wei Jing +1
Multiple-Choice Question Answering (MCQA) is a challenging task in machine reading comprehension. The main challenge in MCQA is to extract "evidence" from the given context that su…
Span-based Localizing Network for Natural Language Video Localization
Hao Zhang, Aixin Sun, Wei Jing +1
Given an untrimmed video and a text query, natural language video localization (NLVL) is to locate a matching span from the video that semantically corresponds to the query. Existi…