8 papers
HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation
Linyin Luo, Yujuan Ding, Yunshan Ma +2
Advanced multimodal Retrieval-Augmented Generation (MRAG) techniques have been widely applied to enhance the capabilities of Large Multimodal Models (LMMs), but they also bring alo…
Topology Sculptor, Shape Refiner: Discrete Diffusion Model for High-Fidelity 3D Meshes Generation
Kaiyu Song, Hanjiang Lai, Yaqing Zhang +3
In this paper, we introduce Topology Sculptor, Shape Refiner (TSSR), a novel method for generating high-quality, artist-style 3D meshes based on Discrete Diffusion Models (DDMs). O…
Flow Matching Posterior Sampling: A Training-free Conditional Generation for Flow Matching
Kaiyu Song, Hanjiang Lai, Yan Pan +2
Training-free conditional generation based on flow matching aims to leverage pre-trained unconditional flow matching models to perform conditional generation without retraining. Re…
Rethinking Oversaturation in Classifier-Free Guidance via Low Frequency
Kaiyu Song, Hanjiang Lai
Classifier-free guidance (CFG) succeeds in condition diffusion models that use a guidance scale to balance the influence of conditional and unconditional terms. A high guidance sca…
Two Simple Principles for Diffusion-Based Test-Time Adaptation
Kaiyu Song, Hanjiang Lai, Yan Pan +2
Recently, diffusion-based test-time adaptations (TTA) have shown great advances, which leverage a diffusion model to map the images in the unknown test domain to the training domai…
Leveraging Previous Steps: A Training-free Fast Solver for Flow Diffusion
Kaiyu Song, Hanjiang Lai
Flow diffusion models (FDMs) have recently shown potential in generation tasks due to the high generation quality. However, the current ordinary differential equation (ODE) solver…