71 citations · 163 across the 26 of their papers we have counts for
45 papers
End-to-end contextual asr based on posterior distribution adaptation for hybrid ctc/attention system
Zhengyi Zhang, Pan Zhou
End-to-end (E2E) speech recognition architectures assemble all components of traditional speech recognition system into a single model. Although it simplifies ASR system, it introd…
Exploring Motion and Appearance Information for Temporal Sentence Grounding
Daizong Liu, Xiaoye Qu, Pan Zhou +1
This paper addresses temporal sentence grounding. Previous works typically solve this task by learning frame-level video features and align them with the textual information. A maj…
Wav-BERT: Cooperative Acoustic and Linguistic Representation Learning for Low-Resource Speech Recognition
Guolin Zheng, Yubei Xiao, Ke Gong +3
Unifying acoustic and linguistic representation learning has become increasingly crucial to transfer the knowledge learned on the abundance of high-resource language data for low-r…
Progressively Guide to Attend: An Iterative Alignment Framework for Temporal Sentence Grounding
Daizong Liu, Xiaoye Qu, Pan Zhou
A key solution to temporal sentence grounding (TSG) exists in how to learn effective alignment between vision and language features extracted from an untrimmed video and a sentence…
Adaptive Proposal Generation Network for Temporal Sentence Localization in Videos
Daizong Liu, Xiaoye Qu, Jianfeng Dong +1
We address the problem of temporal sentence localization in videos (TSLV). Traditional methods follow a top-down framework which localizes the target segment with pre-defined segme…
Coarse to Fine: Domain Adaptive Crowd Counting via Adversarial Scoring Network
Zhikang Zou, Xiaoye Qu, Pan Zhou +4
Recent deep networks have convincingly demonstrated high capability in crowd counting, which is a critical task attracting widespread attention due to its various industrial applic…