collaborators

5 papers

cs.CV2026

V-FIND: Revealing the Intrinsic Forgery Knowledge Encoded in Video Forgery Detectors

Shichao Kan, Chengpeng Hong, Jingtong Dou +8

As generated videos become increasingly realistic, reliable video forgery detection is increasingly important. Existing studies typically optimize and use video forgery detectors a…

cs.CV2026

Not all tokens contribute equally to diffusion learning

Guoqing Zhang, Lu Shi, Wanru Xu +4

With the rapid development of conditional diffusion models, significant progress has been made in text-to-video generation. However, we observe that these models often neglect sema…

cs.CV2025

Condition Weaving Meets Expert Modulation: Towards Universal and Controllable Image Generation

Guoqing Zhang, Xingtong Ge, Lu Shi +5

The image-to-image generation task aims to produce controllable images by leveraging conditional inputs and prompt instructions. However, existing methods often train separate cont…

cs.LG2025

Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental Learning

Haojie Zhang, Yixiong Liang, Hulin Kuang +5

Multimodal Biomedical Image Incremental Learning (MBIIL) is essential for handling diverse tasks and modalities in the biomedical domain, as training separate models for each modal…

cs.CV2025

Object Retrieval for Visual Question Answering with Outside Knowledge

Shichao Kan, Yuhai Deng, Jiale Fu +5

Retrieval-augmented generation (RAG) with large language models (LLMs) plays a crucial role in question answering, as LLMs possess limited knowledge and are not updated with contin…