6 papers
FANeRV: Frequency Separation and Augmentation based Neural Representation for Video
Li Yu, Zhihui Li, Chao Yao +2
Neural representations for video (NeRV) have gained considerable attention for their strong performance across various video tasks. However, existing NeRV methods often struggle to…
Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
Li Yu, Xuanzhe Sun, Wei Zhou +1
Video saliency prediction is crucial for downstream applications, such as video compression and human-computer interaction. With the flourishing of multimodal learning, researchers…
DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment
Li Yu, Situo Wang, Wei Zhou +1
Inspired by the dual-stream theory of the human visual system (HVS) - where the ventral stream is responsible for object recognition and detail analysis, while the dorsal stream fo…
Relevance-guided Audio Visual Fusion for Video Saliency Prediction
Li Yu, Xuanzhe Sun, Pan Gao +1
Audio data, often synchronized with video frames, plays a crucial role in guiding the audience's visual attention. Incorporating audio information into video saliency prediction ta…
Multi-task Feature Enhancement Network for No-Reference Image Quality Assessment
Li Yu
Due to the scarcity of labeled samples in Image Quality Assessment (IQA) datasets, numerous recent studies have proposed multi-task based strategies, which explore feature informat…
High-Frequency Enhanced Hybrid Neural Representation for Video Compression
Li Yu, Zhihui Li, Jimin Xiao +1
Neural Representations for Videos (NeRV) have simplified the video codec process and achieved swift decoding speeds by encoding video content into a neural network, presenting a pr…