28 citations · 92 across the 52 of their papers we have counts for
13 papers · 1 filter
Taming CoT Obfuscation in VLMs: From Mechanistic Evidence to Activation Enforcement
Xutao Mao, Jianing Zhu, Jinman Zhao +4
Reinforcement learning (RL) improves reasoning in vision-language models (VLMs) but can induce chain-of-thought (CoT) obfuscation: an operational, non-intentional outcome where tas…
Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment
Suqin Yuan, Runqi Lin, Muyang Li +5
Human-feedback alignment has made language models useful assistants and is commonly described as aligning them with humans. However, the responses people prefer from an AI need not…
Generated Images Are Easier to Forget: A Machine Unlearning Perspective for Synthetic Image Detection
Jun Nie, Yonggang Zhang, Tongliang Liu +3
Robust detection of generated images is critical to counter the misuse of generative models. Existing methods primarily depend on learning from human-annotated training datasets, l…
MLUBench: A Benchmark for Lifelong Unlearning Evaluation in MLLMs
He Li, Haoang Chi, Qizhou Wang +6
Multimodal large language models (MLLMs) are trained on massive multimodal data, making data unlearning increasingly important as data owners may request the removal of specific co…
AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions
Jingwei Sun, Jianing Zhu, Yuanyi Li +3
Autonomous computer use agents that powered by multimodal large language models (MLLMs) are emerging as capable assistants for completing complex digital workflows. However, real-w…
Rethinking How to Remember: Beyond Atomic Facts in Lifelong LLM Agent Memory
Jingwei Sun, Jianing Zhu, Jiangchao Yao +2
To enable reliable long-term interaction, LLM agents require a memory system that can faithfully store, efficiently retrieve, and deeply reason over accumulated dialogue history. M…