5 papers
LIVEditor-14B: Lightning Unified Video Editing via In-Context Sparse Attention
Shitong Shao, Zikai Zhou, Haopeng Li +4
Video editing has evolved toward In-Context Learning (ICL) paradigms, yet the resulting quadratic attention costs create a critical computational bottleneck. In this work, we propo…
Answering from Sure to Uncertain: Uncertainty-Aware Curriculum Learning for Video Question Answering
Haopeng Li, Mohammed Bennamoun, Jun Liu +2
While significant advancements have been made in video question answering (VideoQA), the potential benefits of enhancing model generalization through tailored difficulty scheduling…
Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional Sports
Haopeng Li, Andong Deng, Jun Liu +5
Reasoning over sports videos for question answering is an important task with numerous applications, such as player training and information retrieval. However, this task has not b…
Admitting Ignorance Helps the Video Question Answering Models to Answer
Haopeng Li, Tom Drummond, Mingming Gong +2
Significant progress has been made in the field of video question answering (VideoQA) thanks to deep learning and large-scale pretraining. Despite the presence of sophisticated mod…
Multimodal Semantic Communication for Generative Audio-Driven Video Conferencing
Haonan Tong, Haopeng Li, Hongyang Du +3
This paper studies an efficient multimodal data communication scheme for video conferencing. In our considered system, a speaker gives a talk to the audiences, with talking head vi…