collaborators

5 papers

cs.CR2026

RAS: Measuring LLM Safety Through Refusal Alignment

Chang-Chieh Huang, Yan-Lun Chen, Chia-Mu Yu +1

Safety evaluation of large language models (LLMs) is commonly performed by querying models with unsafe or jailbreak prompts and judging whether their outputs violate a safety polic…

cs.CR2026

Tracing Target Answers in Poisoned Retrieval Corpora via Token Influence Attribution

Yan-Lun Chen, Pin-Yu Chen, Chia-Mu Yu +3

Retrieval-Augmented Generation (RAG) systems are vulnerable to corpus poisoning attacks that manipulate model outputs through malicious retrieved documents. Existing detection meth…

cs.CR2026

IU: Imperceptible Universal Backdoor Attack

Hsin Lin, Yan-Lun Chen, Ren-Hung Hwang +1

Backdoor attacks pose a critical threat to the security of deep neural networks, yet existing efforts on universal backdoors often rely on visually salient patterns, making them ea…

cs.LG2025

BADTV: Unveiling Backdoor Threats in Third-Party Task Vectors

Chia-Yi Hsu, Yu-Lin Tsai, Yu Zhe +6

Task arithmetic in large-scale pre-trained models enables agile adaptation to diverse downstream tasks without extensive retraining. By leveraging task vectors (TVs), users can per…

cs.CL2025

Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge

Yan-Lun Chen, Yi-Ru Wei, Chia-Yi Hsu +5

Large language models (LLMs) demonstrate strong task-specific capabilities through fine-tuning, but merging multiple fine-tuned models often leads to degraded performance due to ov…