Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
Kaisi Guan, Xihua Wang, Zhengfeng Lai +5
This study focuses on a challenging yet promising task, Text-to-Sounding-Video (T2SV) generation, which aims to generate a video with synchronized audio from text conditions, meanw…
cs.CV2026
Incentivizing Temporal-Awareness in Egocentric Video Understanding Models
Zhiyang Xu, Tian Qin, Bowen Jin +4
Multimodal large language models (MLLMs) have recently shown strong performance in visual understanding, yet they often lack temporal awareness, particularly in egocentric settings…