collaborators

5 papers

cs.CV2026

ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation

Linhan Cao, Siyuan Li, Jun Lan +8

Large multimodal models (LMMs) have demonstrated strong OCR recognition capabilities, yet remain vulnerable to adversarial visual text that is readable to humans but challenging fo…

cs.AI2026

RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection

Yan Hong, Wei Li, Kedong Xiu +6

Knowledge injection updates pretrained MLLMs with new factual or domain-specific knowledge, but fitting full authoritative answers can cause drift in non-updated behavior. Online d…

stat.ML2026

LoMC: Localized Multidirectional Correction for Refusal Suppression in Routed Foundation Models

Yan Hong, Kedong Xiu, Wei Li +6

We study controlled post-training refusal suppression in routed MoE and hybrid-MoE foundation models, aiming to increase non-refusal target-response behavior while preserving gener…

cs.CR2025

SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning

Kaiwen Zhou, Ahmed Elgohary, A S M Iftekhar +1

The ability of LLM agents to plan and invoke tools exposes them to new safety risks, making a comprehensive red-teaming system crucial for discovering vulnerabilities and ensuring…

cs.CR2025

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models

Yanjiang Liu, Shuhen Zhou, Yaojie Lu +6

Automated red-teaming has become a crucial approach for uncovering vulnerabilities in large language models (LLMs). However, most existing methods focus on isolated safety flaws, l…