papers

Publications (14)

cs.CV2026

Training-free Motion Factorization for Compositional Video Generation

Zixuan Wang, Ziqin Zhou, Feng Chen +4

Compositional video generation aims to synthesize multiple instances with diverse appearance and motion. However, current approaches mainly focus on binding semantics, neglecting t…

cs.CV2026

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding

Haiyue Zhang, Yi Bin, Xun Jiang +5

VisualRouter is a training-free, plug‑and‑play framework that classifies queries as global or local and applies tailored visual sampling strategies to select informative frames, im…

#video understanding#visual sampling#query grounding#long videos
cs.CV2025

Unified Prompt Attack Against Text-to-Image Generation Models

Duo Peng, Qiuhong Ke, Mark He Huang +2

Text-to-Image (T2I) models have advanced significantly, but their growing popularity raises security concerns due to their potential to generate harmful images. To address these is…

cs.CV2024

UPAM: Unified Prompt Attack in Text-to-Image Generation Models Against Both Textual Filters and Visual Checkers

Duo Peng, Qiuhong Ke, Jun Liu

Text-to-Image (T2I) models have raised security concerns due to their potential to generate inappropriate or harmful images. In this paper, we propose UPAM, a novel framework that…

cs.CV2023

Unsupervised Domain Adaptation via Domain-Adaptive Diffusion

Duo Peng, Qiuhong Ke, Yinjie Lei +1

Unsupervised Domain Adaptation (UDA) is quite challenging due to the large distribution discrepancy between the source domain and the target domain. Inspired by diffusion models wh…

cs.LG2026

PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR

James Burgess, Jan N. Hansen, Duo Peng +5

Search agents are language models (LMs) that reason and search knowledge bases (or the web) to answer questions; recent methods supervise only the final answer accuracy using reinf…

cs.CV2023

Diffusion-based Image Translation with Label Guidance for Domain Adaptive Semantic Segmentation

Duo Peng, Ping Hu, Qiuhong Ke +1

Translating images from a source domain to a target domain for learning target models is one of the most common strategies in domain adaptive semantic segmentation (DASS). However,…

cs.CV2021

Global and Local Texture Randomization for Synthetic-to-Real Semantic Segmentation

Duo Peng, Yinjie Lei, Lingqiao Liu +2

Semantic segmentation is a crucial image understanding task, where each pixel of image is categorized into a corresponding label. Since the pixel-wise labeling for ground-truth is…

cs.CV2020

Hierarchical Paired Channel Fusion Network for Street Scene Change Detection

Yinjie Lei, Duo Peng, Pingping Zhang +2

Street Scene Change Detection (SSCD) aims to locate the changed regions between a given street-view image pair captured at different times, which is an important yet challenging ta…

cs.CV2022

Semantic-Aware Domain Generalized Segmentation

Duo Peng, Yinjie Lei, Munawar Hayat +2

Deep models trained on source domain lack generalization when evaluated on unseen target domains with different data distributions. The problem becomes even more pronounced when we…

cs.CV2025

Visual Prompting for One-shot Controllable Video Editing without Inversion

Zhengbo Zhang, Yuxi Zhou, Duo Peng +4

One-shot controllable video editing (OCVE) is an important yet challenging task, aiming to propagate user edits that are made -- using any image editing tool -- on the first frame…

cs.CV2025

Training-free Dense-Aligned Diffusion Guidance for Modular Conditional Image Synthesis

Zixuan Wang, Duo Peng, Feng Chen +2

Conditional image synthesis is a crucial task with broad applications, such as artistic creation and virtual reality. However, current generative methods are often task-oriented wi…

cs.CV2024

Diff-Tracker: Text-to-Image Diffusion Models are Unsupervised Trackers

Zhengbo Zhang, Li Xu, Duo Peng +2

We introduce Diff-Tracker, a novel approach for the challenging unsupervised visual tracking task leveraging the pre-trained text-to-image diffusion model. Our main idea is to leve…

cs.CV2021

Sparse-to-dense Feature Matching: Intra and Inter domain Cross-modal Learning in Domain Adaptation for 3D Semantic Segmentation

Duo Peng, Yinjie Lei, Wen Li +2

Domain adaptation is critical for success when confronting with the lack of annotations in a new domain. As the huge time consumption of labeling process on 3D point cloud, domain…