28 papers
Targeted Counterfactual Fingerprinting for Black-Box LLM Ownership Verification
Yutong Wu, Xiaofan Bai, Shixin Li +10
Large language models (LLMs) are high-value assets that can be derived through redeployment, fine-tuning, quantization, or further alignment. Because deployed LLMs are commonly exp…
TYPO: Instruction-Dense Visual Jailbreaks against Commercial Closed-Source Image-Generation Models
Meng Xie, Li Zeng, Hangtao Zhang +4
Recent commercial image-generation models can generate high-quality images with readable text (e.g., posters, infographics, and manuals), attracting considerable attention. Yet we…
When Does Muon Help Agentic Reinforcement Learning?
Kai Ruan, Jinghao Lin, Zihe Huang +4
Muon is competitive with AdamW in large-scale pre-training, but its operating regime in reinforcement-learning post-training remains unclear. We map this regime on ALFWorld, a spar…
Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade
Kai Ruan, Zihe Huang, Ziqi Zhou +4
The paper proposes using lightweight linear probes on hidden states of large language model agents to predict failures early and abort doomed episodes, achieving large compute savi…
Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection
Qi Lu, Ziqi Zhou, Yufei Song +5
Vision Language Models (VLMs) offer powerful multimodal ability but also expose users to text-based privacy attacks where adversaries crawl online photos and query VLMs to extract…
Dual-branch Robust Unlearnable Examples
Xianlong Wang, Hangtao Zhang, Wenbo Pan +4
Unlearnable examples (UEs) aim to compromise model training by injecting imperceptible perturbations to clean samples. However, existing UE schemes exhibit limited robustness again…