Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
CodecCap: High-Fidelity Codec-Inspired Residual Modeling for Dense Video Captioning
Zihan Lin, Songhe Deng, Shuwei He +6
Existing video captioning methods struggle to balance visual fidelity and redundancy: holistic captions are compact but lose fine-grained evidence, whereas segment-wise captions im…
cs.CV2025
Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
Rui Liu, Shuwei He, Yifan Hu +1
Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize the reverberant speech for the spoken content. The challenge of this task lies in unde…