activity
20242026
collaborators

11 papers

cs.CL2026

Hey, That's My Data! Token-Only Dataset Inference in Large Language Models

Chen Xiong, Zihao Wang, Rui Zhu +4

Large Language Models (LLMs) rely on massive training datasets, often including proprietary data, which raises concerns about unauthorized usage and copyright infringement. Existin…

cs.LG2026

Defining and Evaluating Physical Safety for Large Language Models

Yung-Chen Tang, Pin-Yu Chen, Tsung-Yi Ho

Large Language Models (LLMs) are increasingly used to control robotic systems such as drones, but their risks of causing physical threats and harm in real-world applications remain…

cs.AI2025

CoP: Agentic Red-teaming for Large Language Models using Composition of Principles

Chen Xiong, Pin-Yu Chen, Tsung-Yi Ho

Recent advances in Large Language Models (LLMs) have spurred transformative applications in various domains, ranging from open-source to proprietary LLMs. However, jailbreak attack…

cs.CR2025

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs

Xiaomeng Hu, Pin-Yu Chen, Tsung-Yi Ho

As large language models (LLMs) become more integral to society and technology, ensuring their safety becomes essential. Jailbreak attacks exploit vulnerabilities to bypass safety…

cs.MM2025

Optimization-Free Universal Watermark Forgery with Regenerative Diffusion Models

Chaoyi Zhu, Zaitang Li, Renyi Yang +4

Watermarking becomes one of the pivotal solutions to trace and verify the origin of synthetic images generated by artificial intelligence models, but it is not free of risks. Recen…

cs.CR2025

Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets

Lei Hsiung, Tianyu Pang, Yung-Chen Tang +4

Recent advancements in large language models (LLMs) have underscored their vulnerability to safety alignment jailbreaks, particularly when subjected to downstream fine-tuning. Howe…