works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.CV2026

PersGuard: Preventing Malicious Personalization in Text-to-Image Diffusion Models via Model Backdoors

Xinwei Liu, Xiaojun Jia, Yuan Xun +2

The paper proposes PersGuard, a backdoor-based method that embeds protective triggers into pre‑trained text‑to‑image diffusion models so that unauthorized fine‑tuning on protected…

cs.CV2026

Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation

Ruoyu Chen, Xiaoqing Guo, Kangwei Liu +6

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in aligning visual inputs with natural language outputs. Yet, the extent to which generated token…

cs.CV2025

FaceInsight: A Multimodal Large Language Model for Face Perception

Jingzhi Li, Changjiang Luo, Ruoyu Chen +4

Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perfo…

cs.LG2025

Beyond Progress Measures: Theoretical Insights into the Mechanism of Grokking

Zihan Gu, Ruoyu Chen, Hua Zhang +2

Grokking, referring to the abrupt improvement in test accuracy after extended overfitting, offers valuable insights into the mechanisms of model generalization. Existing researches…

cs.CV2025

Generalized Semantic Contrastive Learning via Embedding Side Information for Few-Shot Object Detection

Ruoyu Chen, Hua Zhang, Jingzhi Li +3

The objective of few-shot object detection (FSOD) is to detect novel objects with few training samples. The core challenge of this task is how to construct a generalized feature sp…

cs.CV2025

Interpreting Object-level Foundation Models via Visual Precision Search

Ruoyu Chen, Siyuan Liang, Jingzhi Li +5

Advances in multimodal pre-training have propelled object-level foundation models, such as Grounding DINO and Florence-2, in tasks like visual grounding and object detection. Howev…