activity
20232025
most citedEfficient Multimodal Fusion via Interactive Prompting

5 citations · 6 across the 9 of their papers we have counts for

collaborators

9 papers

cs.CV2025

Adversarial-Guided Diffusion for Multimodal LLM Attacks

Chengwei Xia, Fan Ma, Ruijie Quan +2

This paper addresses the challenge of generating adversarial image using a diffusion model to deceive multimodal large language models (MLLMs) into generating the targeted response…

cs.CV20241 cited

Autonomous LLM-Enhanced Adversarial Attack for Text-to-Motion

Honglei Miao, Fan Ma, Ruijie Quan +2

Human motion generation driven by deep generative models has enabled compelling applications, but the ability of text-to-motion (T2M) models to produce realistic motions from text…

cs.CV2024

Shape2Scene: 3D Scene Representation Learning Through Pre-training on Shape Data

Tuo Feng, Wenguan Wang, Ruijie Quan +1

Current 3D self-supervised learning methods of 3D scenes face a data desert issue, resulting from the time-consuming and expensive collecting process of 3D scene data. Conversely,…

cs.CV2024

General and Task-Oriented Video Segmentation

Mu Chen, Liulei Li, Wenguan Wang +2

We present GvSeg, a general video segmentation framework for addressing four different video segmentation tasks (i.e., instance, semantic, panoptic, and exemplar-guided) while main…

cs.CV2024

AudioScenic: Audio-Driven Video Scene Editing

Kaixin Shen, Ruijie Quan, Linchao Zhu +2

Audio-driven visual scene editing endeavors to manipulate the visual background while leaving the foreground content unchanged, according to the given audio signals. Unlike current…

cs.AI2024

Neural Interaction Energy for Multi-Agent Trajectory Prediction

Kaixin Shen, Ruijie Quan, Linchao Zhu +2

Maintaining temporal stability is crucial in multi-agent trajectory prediction. Insufficient regularization to uphold this stability often results in fluctuations in kinematic stat…