5 papers · 1 filter
LIVEditor-14B: Lightning Unified Video Editing via In-Context Sparse Attention
Shitong Shao, Zikai Zhou, Haopeng Li +4
Video editing has evolved toward In-Context Learning (ICL) paradigms, yet the resulting quadratic attention costs create a critical computational bottleneck. In this work, we propo…
Admitting Ignorance Helps the Video Question Answering Models to Answer
Haopeng Li, Tom Drummond, Mingming Gong +2
Significant progress has been made in the field of video question answering (VideoQA) thanks to deep learning and large-scale pretraining. Despite the presence of sophisticated mod…
Answering from Sure to Uncertain: Uncertainty-Aware Curriculum Learning for Video Question Answering
Haopeng Li, Mohammed Bennamoun, Jun Liu +2
While significant advancements have been made in video question answering (VideoQA), the potential benefits of enhancing model generalization through tailored difficulty scheduling…
Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional Sports
Haopeng Li, Andong Deng, Jun Liu +5
Reasoning over sports videos for question answering is an important task with numerous applications, such as player training and information retrieval. However, this task has not b…
Reconstructive Sequence-Graph Network for Video Summarization
Bin Zhao, Haopeng Li, Xiaoqiang Lu +1
Exploiting the inner-shot and inter-shot dependencies is essential for key-shot based video summarization. Current approaches mainly devote to modeling the video as a frame sequenc…