activity
20232026
most citedTowards Explainable Fake Image Detection with Multi-Modal Large Language Models

2 citations · 4 across the 21 of their papers we have counts for

collaborators

23 papers

cs.AI2026

Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents

Zhaoyuan Huang, Tianjie Ju, Pengzhou Cheng +7

Graphical user interface (GUI) agents are increasingly used to execute natural-language instructions on user interfaces, yet real users may issue infeasible instructions due to ben…

cs.CV2026

ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation

Linhan Cao, Siyuan Li, Jun Lan +8

Large multimodal models (LMMs) have demonstrated strong OCR recognition capabilities, yet remain vulnerable to adversarial visual text that is readable to humans but challenging fo…

cs.CV2026

Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection

Hao Tan, Jun Lan, Zichang Tan +7

The growing capability of image generation models has made synthetic images a routine presence in open media, making robust and generalizable AI-Generated Image (AIGI) detection in…

stat.ML2026

LoMC: Localized Multidirectional Correction for Refusal Suppression in Routed Foundation Models

Yan Hong, Kedong Xiu, Wei Li +6

We study controlled post-training refusal suppression in routed MoE and hybrid-MoE foundation models, aiming to increase non-refusal target-response behavior while preserving gener…

cs.CV2026

Adaptive and Balanced Re-initialization for Long-timescale Continual Test-time Domain Adaptation

Yanshuo Wang, Jinguang Tong, Jun Lan +5

Continual test-time domain adaptation (CTTA) aims to adjust models so that they can perform well over time across non-stationary environments. While previous methods have made cons…

cs.CV2025

Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images

Yikun Ji, Yan Hong, Bowen Deng +5

The rapid growth of AI-generated imagery has blurred the boundary between real and synthetic content, raising practical concerns for digital integrity. Vision-language models (VLMs…