6 papers
NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration: Methods and Results
Wenbin Zou, Tianyi Liu, Kejun Wu +37
This paper reports on the NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration (BSCVR). The challenge aims to advance research on recovering visually coherent videos from…
Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
Junlong Li, Huaiyuan Xu, Sijie Cheng +4
Driven by recent advances in vision-language models (VLMs) and egocentric perception research, the emerging topic of an egocentric procedural AI assistant (EgoProceAssist) is intro…
MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
Qihao Zhao, Yunqi Cao, Yangyu Huang +4
Despite recent advances in multimodal large language models (MLLMs), their ability to understand and interact with music remains limited. Music understanding requires grounded reas…
Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
Tianyi Liu, Kejun Wu, Chen Cai +3
Video signals are vulnerable in multimedia communication and storage systems, as even slight bitstream-domain corruption can lead to significant pixel-domain degradation. To recove…
PromptSR: Cascade Prompting for Lightweight Image Super-Resolution
Wenyang Liu, Chen Cai, Jianjun Gao +4
Although the lightweight Vision Transformer has significantly advanced image super-resolution (SR), it faces the inherent challenge of a limited receptive field due to the window-b…
ByteNet: Rethinking Multimedia File Fragment Classification through Visual Perspectives
Wenyang Liu, Kejun Wu, Tianyi Liu +3
Multimedia file fragment classification (MFFC) aims to identify file fragment types, e.g., image/video, audio, and text without system metadata. It is of vital importance in multim…