Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Unsupervised Collaborative Domain Adaptation for Driving Scene Parsing
Jiahe Fan, Shaolong Shu, Mingjian Sun +4
Reliable driving scene parsing is a fundamental capability for autonomous vehicles operating in open and dynamic driving environments. However, adapting perception models to new de…
cs.CV2025
Contextual AD Narration with Interleaved Multimodal Sequence
Hanlin Wang, Zhan Tong, Kecheng Zheng +2
The Audio Description (AD) task aims to generate descriptions of visual elements for visually impaired individuals to help them access long-form video content, like movies. With vi…
cs.CV2025
Learning Human Skill Generators at Key-Step Levels
Yilu Wu, Chenhui Zhu, Shuai Wang +4
We are committed to learning human skill generators at key-step levels. The generation of skills is a challenging endeavor, but its successful implementation could greatly facilita…