1 citations · 1 across the 9 of their papers we have counts for
12 papers
ARENA: Automated Red-Teaming for Large Audio Language Models
Jiaming He, Zhicong Huang, Tian Jin +5
Large audio-language models (LALMs) make it possible to interact with language models through speech, music, and environmental sound, but they also introduce a safety surface that…
Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery Robustness
Haoting Qian, Qingjie Zhang, Zhicong Huang +2
Benchmarking machine unlearning methods is critical to understand whether sensitive knowledge is removed from large language models (LLMs) or not. Current unlearning benchmarks inc…
DeepInvert: Semi-Supervised Embedding Inversion Against Obfuscated Language Models
Zhicong Huang, Cheng Hong, Tao Wei
Cloud-based language model services routinely process prompts containing sensitive information. Obfuscation-based defenses---including ObfusLM, SentinelLMs, TextObfuscator, and DPN…
TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment
Changyue Li, Jiaming He, Youliang Yuan +4
Fine-Tuning-as-a-Service (FTaaS) platforms let users train large language models (LLMs) on customized tasks, but this pipeline could erode models' safety alignment. In practice, se…
When New Generators Arrive: Lifelong Machine-Generated Text Attribution via Ridge Feature Transfer
Zhen Sun, Yifan Liao, Zhicong Huang +4
Machine-generated text (MGT) attribution aims to identify the specific generator responsible for a given text, thereby providing fine-grained evidence for model accountability and…
Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models
Jialin Wu, Wei Shi, Han Shen +5
Despite the advanced capabilities of Large Vision-Language Models (LVLMs), they frequently suffer from object hallucination. One reason is that visual features and pretrained textu…