3 papers
cs.CV2025
DeCo-VAE: Learning Compact Latents for Video Reconstruction via Decoupled Representation
Xiangchen Yin, Jiahui Yuan, Zhangchi Hu +5
Existing video Variational Autoencoders (VAEs) generally overlook the similarity between frame contents, leading to redundant latent modeling. In this paper, we propose decoupled V…
cs.CV2025
Class-Aware Prototype Learning with Negative Contrast for Test-Time Adaptation of Vision-Language Models
Xiaozhen Qiao, Jingkai Zhao, Yuqiu Jiang +4
Vision-Language Models (VLMs) demonstrate impressive zero-shot generalization through large-scale image-text pretraining, yet their performance can drop once the deployment distrib…
cs.CV2025
Bidirectional Prototype-Reward co-Evolution for Test-Time Adaptation of Vision-Language Models
Xiaozhen Qiao, Peng Huang, Jiakang Yuan +6
Test-time adaptation (TTA) is crucial in maintaining performance of Vision Language Models (VLMs) when facing distribution shifts, particularly when the source data or target label…