14 papers
RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editing
Ruiliang Gong, Zhen Wang, Yanghao Wang +1
Inversion-free flow-based image editing avoids latent inversion, but still requires a target-side state at every editing step. The widely used equal-displacement construction keeps…
h-Flow: Flexible Flow-based Image Editing via Doob's h-Transform
Zehui Guo, Zhen Wang, Junwei Shu +3
Editing images with pre-trained text-to-image flow models typically requires carefully balancing target alignment with the desired prompt and source consistency with the original i…
Robustifying Vision-Language Models via Test-Time Prompt Adaptation
Xingyu Zhu, Huanshen Wu, Shuo Wang +4
Pre-trained Vision-Language Models (VLMs) such as CLIP achieve strong zero-shot generalization, but their performance degrades sharply under adversarial perturbations. Existing tes…
LISA: Likelihood Score Alignment for Visual-condition Controllable Generation
Yanghao Wang, Hongxu Chen, Jiazhen Liu +4
The prevalent dual-branch paradigm, i.e., training a side network to encode visual conditions and fusing its intermediate-layer features to a frozen pretrained main network, has sh…
Bi-Anchor Interpolation Solver for Accelerating Generative Modeling
Hongxu Chen, Hongxiang Li, Zhen Wang +1
Flow Matching (FM) models have emerged as a leading paradigm for high-fidelity synthesis. However, their reliance on iterative Ordinary Differential Equation (ODE) solving creates…
Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment
Cong Wang, Hanxin Zhu, Jiayi Luo +6
Large-scale video generation models have made remarkable progress in semantic consistency and visual quality, producing videos that are increasingly coherent and visually convincin…