activity
20242026
collaborators

28 papers

cs.CR2026

Targeted Counterfactual Fingerprinting for Black-Box LLM Ownership Verification

Yutong Wu, Xiaofan Bai, Shixin Li +10

Large language models (LLMs) are high-value assets that can be derived through redeployment, fine-tuning, quantization, or further alignment. Because deployed LLMs are commonly exp…

cs.CR2026

TYPO: Instruction-Dense Visual Jailbreaks against Commercial Closed-Source Image-Generation Models

Meng Xie, Li Zeng, Hangtao Zhang +4

Recent commercial image-generation models can generate high-quality images with readable text (e.g., posters, infographics, and manuals), attracting considerable attention. Yet we…

cs.LG2026

When Does Muon Help Agentic Reinforcement Learning?

Kai Ruan, Jinghao Lin, Zihe Huang +4

Muon is competitive with AdamW in large-scale pre-training, but its operating regime in reinforcement-learning post-training remains unclear. We map this regime on ALFWorld, a spar…

cs.AI2026

Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade

Kai Ruan, Zihe Huang, Ziqi Zhou +4

The paper proposes using lightweight linear probes on hidden states of large language model agents to predict failures early and abort doomed episodes, achieving large compute savi…

cs.CV2026

Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection

Qi Lu, Ziqi Zhou, Yufei Song +5

Vision Language Models (VLMs) offer powerful multimodal ability but also expose users to text-based privacy attacks where adversaries crawl online photos and query VLMs to extract…

cs.CV2026

Dual-branch Robust Unlearnable Examples

Xianlong Wang, Hangtao Zhang, Wenbo Pan +4

Unlearnable examples (UEs) aim to compromise model training by injecting imperceptible perturbations to clean samples. However, existing UE schemes exhibit limited robustness again…