activity
20242026
collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

Olaf-World: Orienting Latent Actions for Video World Modeling

Yuxin Jiang, Yuchao Gu, Ivor W. Tsang +1

Scaling action-controllable world models is limited by the scarcity of action labels. While latent action learning promises to extract control interfaces from unlabeled video, lear…

cs.CV2026

Catch Me If You Can Describe Me: Open-Vocabulary Camouflaged Instance Segmentation with Diffusion

Tuan-Anh Vu, Duc Thanh Nguyen, Qing Guo +4

Text-to-image diffusion techniques have shown exceptional capabilities in producing high-quality, dense visual predictions from open-vocabulary text. This indicates a strong correl…

cs.CV2025

Decouple before Align: Visual Disentanglement Enhances Prompt Tuning

Fei Zhang, Tianfei Zhou, Jiangchao Yao +3

Prompt tuning (PT), as an emerging resource-efficient fine-tuning paradigm, has showcased remarkable effectiveness in improving the task-specific transferability of vision-language…

cs.CV2025

Multi-Modal Dataset Distillation in the Wild

Zhuohang Dang, Minnan Luo, Chengyou Jia +3

Recent multi-modal models have shown remarkable versatility in real-world applications. However, their rapid development encounters two critical data challenges. First, the trainin…

cs.CV2025

Training-Free Dataset Pruning for Instance Segmentation

Yalun Dai, Lingao Xiao, Ivor W. Tsang +1

Existing dataset pruning techniques primarily focus on classification tasks, limiting their applicability to more complex and practical tasks like instance segmentation. Instance s…

cs.CV2024

Multisize Dataset Condensation

Yang He, Lingao Xiao, Joey Tianyi Zhou +1

While dataset condensation effectively enhances training efficiency, its application in on-device scenarios brings unique challenges. 1) Due to the fluctuating computational resour…