Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
Li Yu, Xuanzhe Sun, Wei Zhou +1
Video saliency prediction is crucial for downstream applications, such as video compression and human-computer interaction. With the flourishing of multimodal learning, researchers…
cs.CV2025
DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment
Li Yu, Situo Wang, Wei Zhou +1
Inspired by the dual-stream theory of the human visual system (HVS) - where the ventral stream is responsible for object recognition and detail analysis, while the dorsal stream fo…