13 papers
Hey, That's My Data! Token-Only Dataset Inference in Large Language Models
Chen Xiong, Zihao Wang, Rui Zhu +4
Large Language Models (LLMs) rely on massive training datasets, often including proprietary data, which raises concerns about unauthorized usage and copyright infringement. Existin…
OrgAgent: Organize Your Multi-Agent System like a Company
Yiru Wang, Xinyue Shen, Yaohui Han +3
While large language model-based multi-agent systems have shown strong potential for complex reasoning, how to effectively organize multiple agents remains an open question. In thi…
Defining and Evaluating Physical Safety for Large Language Models
Yung-Chen Tang, Pin-Yu Chen, Tsung-Yi Ho
Large Language Models (LLMs) are increasingly used to control robotic systems such as drones, but their risks of causing physical threats and harm in real-world applications remain…
Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models
Chen Xiong, Zhiyuan He, Pin-Yu Chen +2
Activation steering is a practical post-training model alignment technique to enhance the utility of Large Language Models (LLMs). Prior to deploying a model as a service, develope…
CoP: Agentic Red-teaming for Large Language Models using Composition of Principles
Chen Xiong, Pin-Yu Chen, Tsung-Yi Ho
Recent advances in Large Language Models (LLMs) have spurred transformative applications in various domains, ranging from open-source to proprietary LLMs. However, jailbreak attack…
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
Xiaomeng Hu, Pin-Yu Chen, Tsung-Yi Ho
As large language models (LLMs) become more integral to society and technology, ensuring their safety becomes essential. Jailbreak attacks exploit vulnerabilities to bypass safety…