activity
20242026
collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2026

In-Loop Model Adaptation with Coupled Latent-Noise Guidance for High-Fidelity Subject-Driven Text-to-Image Generation

Yushun Tang, Weiming Chen, Siyi Liu +3

Text-to-image diffusion models have achieved remarkable success in generating high-quality images from a given text prompt. Subject-driven generation aims to synthesize customized…

cs.CV2025

Progressive Conditioned Scale-Shift Recalibration of Self-Attention for Online Test-time Adaptation

Yushun Tang, Ziqiong Liu, Jiyuan Jia +2

Online test-time adaptation aims to dynamically adjust a network model in real-time based on sequential input samples during the inference stage. In this work, we find that, when a…

cs.CV2024

Domain-Conditioned Transformer for Fully Test-time Adaptation

Yushun Tang, Shuoshuo Chen, Jiyuan Jia +2

Fully test-time adaptation aims to adapt a network model online based on sequential analysis of input samples during the inference stage. We observe that, when applying a transform…

cs.CV2024

Learning Visual Conditioning Tokens to Correct Domain Shift for Fully Test-time Adaptation

Yushun Tang, Shuoshuo Chen, Zhehan Kan +3

Fully test-time adaptation aims to adapt the network model based on sequential analysis of input samples during the inference stage to address the cross-domain performance degradat…

cs.CV2024

Conceptual Codebook Learning for Vision-Language Models

Yi Zhang, Ke Yu, Siqi Wu +1

In this paper, we propose Conceptual Codebook Learning (CoCoLe), a novel fine-tuning method for vision-language models (VLMs) to address the challenge of improving the generalizati…

cs.CV2024

NODE-Adapter: Neural Ordinary Differential Equations for Better Vision-Language Reasoning

Yi Zhang, Chun-Wun Cheng, Ke Yu +3

In this paper, we consider the problem of prototype-based vision-language reasoning problem. We observe that existing methods encounter three major challenges: 1) escalating resour…