2 citations · 4 across the 21 of their papers we have counts for
23 papers
Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents
Zhaoyuan Huang, Tianjie Ju, Pengzhou Cheng +7
Graphical user interface (GUI) agents are increasingly used to execute natural-language instructions on user interfaces, yet real users may issue infeasible instructions due to ben…
ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation
Linhan Cao, Siyuan Li, Jun Lan +8
Large multimodal models (LMMs) have demonstrated strong OCR recognition capabilities, yet remain vulnerable to adversarial visual text that is readable to humans but challenging fo…
Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection
Hao Tan, Jun Lan, Zichang Tan +7
The growing capability of image generation models has made synthetic images a routine presence in open media, making robust and generalizable AI-Generated Image (AIGI) detection in…
LoMC: Localized Multidirectional Correction for Refusal Suppression in Routed Foundation Models
Yan Hong, Kedong Xiu, Wei Li +6
We study controlled post-training refusal suppression in routed MoE and hybrid-MoE foundation models, aiming to increase non-refusal target-response behavior while preserving gener…
Adaptive and Balanced Re-initialization for Long-timescale Continual Test-time Domain Adaptation
Yanshuo Wang, Jinguang Tong, Jun Lan +5
Continual test-time domain adaptation (CTTA) aims to adjust models so that they can perform well over time across non-stationary environments. While previous methods have made cons…
Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images
Yikun Ji, Yan Hong, Bowen Deng +5
The rapid growth of AI-generated imagery has blurred the boundary between real and synthetic content, raising practical concerns for digital integrity. Vision-language models (VLMs…