collaborators

8 papers

cs.CV2025

HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation

Linyin Luo, Yujuan Ding, Yunshan Ma +2

Advanced multimodal Retrieval-Augmented Generation (MRAG) techniques have been widely applied to enhance the capabilities of Large Multimodal Models (LMMs), but they also bring alo…

cs.CV2025

Topology Sculptor, Shape Refiner: Discrete Diffusion Model for High-Fidelity 3D Meshes Generation

Kaiyu Song, Hanjiang Lai, Yaqing Zhang +3

In this paper, we introduce Topology Sculptor, Shape Refiner (TSSR), a novel method for generating high-quality, artist-style 3D meshes based on Discrete Diffusion Models (DDMs). O…

cs.CV2025

Flow Matching Posterior Sampling: A Training-free Conditional Generation for Flow Matching

Kaiyu Song, Hanjiang Lai, Yan Pan +2

Training-free conditional generation based on flow matching aims to leverage pre-trained unconditional flow matching models to perform conditional generation without retraining. Re…

cs.CV2025

Rethinking Oversaturation in Classifier-Free Guidance via Low Frequency

Kaiyu Song, Hanjiang Lai

Classifier-free guidance (CFG) succeeds in condition diffusion models that use a guidance scale to balance the influence of conditional and unconditional terms. A high guidance sca…

cs.LG2025

Two Simple Principles for Diffusion-Based Test-Time Adaptation

Kaiyu Song, Hanjiang Lai, Yan Pan +2

Recently, diffusion-based test-time adaptations (TTA) have shown great advances, which leverage a diffusion model to map the images in the unknown test domain to the training domai…

cs.CV2025

Leveraging Previous Steps: A Training-free Fast Solver for Flow Diffusion

Kaiyu Song, Hanjiang Lai

Flow diffusion models (FDMs) have recently shown potential in generation tasks due to the high generation quality. However, the current ordinary differential equation (ODE) solver…