Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Calibrated Multimodal Representation Learning with Missing Modalities
Xiaohao Liu, Xiaobo Xia, Jiaheng Wei +4
Multimodal representation learning harmonizes distinct modalities by aligning them into a unified latent space. Recent research generalizes traditional cross-modal alignment to pro…
cs.CV2025
UtilGen: Utility-Centric Generative Data Augmentation with Dual-Level Task Adaptation
Jiyu Guo, Shuo Yang, Yiming Huang +6
Data augmentation using generative models has emerged as a powerful paradigm for enhancing performance in computer vision tasks. However, most existing augmentation approaches prim…
cs.CV2025
Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models
Yuchen Zhou, Jiayu Tang, Shuo Yang +6
Vision-Language Models (VLMs), exemplified by CLIP, have emerged as foundational for multimodal intelligence. However, their capacity for logical understanding remains significantl…