10 papers · 1 filter
In-Loop Model Adaptation with Coupled Latent-Noise Guidance for High-Fidelity Subject-Driven Text-to-Image Generation
Yushun Tang, Weiming Chen, Siyi Liu +3
Text-to-image diffusion models have achieved remarkable success in generating high-quality images from a given text prompt. Subject-driven generation aims to synthesize customized…
Progressive Conditioned Scale-Shift Recalibration of Self-Attention for Online Test-time Adaptation
Yushun Tang, Ziqiong Liu, Jiyuan Jia +2
Online test-time adaptation aims to dynamically adjust a network model in real-time based on sequential input samples during the inference stage. In this work, we find that, when a…
Domain-Conditioned Transformer for Fully Test-time Adaptation
Yushun Tang, Shuoshuo Chen, Jiyuan Jia +2
Fully test-time adaptation aims to adapt a network model online based on sequential analysis of input samples during the inference stage. We observe that, when applying a transform…
Learning Visual Conditioning Tokens to Correct Domain Shift for Fully Test-time Adaptation
Yushun Tang, Shuoshuo Chen, Zhehan Kan +3
Fully test-time adaptation aims to adapt the network model based on sequential analysis of input samples during the inference stage to address the cross-domain performance degradat…
Conceptual Codebook Learning for Vision-Language Models
Yi Zhang, Ke Yu, Siqi Wu +1
In this paper, we propose Conceptual Codebook Learning (CoCoLe), a novel fine-tuning method for vision-language models (VLMs) to address the challenge of improving the generalizati…
NODE-Adapter: Neural Ordinary Differential Equations for Better Vision-Language Reasoning
Yi Zhang, Chun-Wun Cheng, Ke Yu +3
In this paper, we consider the problem of prototype-based vision-language reasoning problem. We observe that existing methods encounter three major challenges: 1) escalating resour…