15 papers
Poison with Style: A Practical Poisoning Attack on Code Large Language Models
Khang Tran, Yazan Boshmaf, Issa Khalil +3
Code Large Language Models (CLLMs) serve as the core of modern code agents, enabling developers to automate complex software development tasks. In this paper, we present Poison-wit…
Gradient Transformer: Learning to Generate Updates for LLMs
Binh-Nguyen Nguyen, Khang Tran, NhatHai Phan +1
Many organizations lack computational resources to fine-tune large language models (LLMs) on private (unshareable) data for better utility, while fine-tuning tiny language models (…
Robust LLM Watermarking with Minimal Semantic Distortion for IP Protection
Kieu Dang, Phung Lai, NhatHai Phan +2
Proprietary large language models (LLMs) face risks of intellectual property (IP) violation, as adversaries can replicate an LLM by collecting input-output pairs to train a surroga…
PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection
Tuan Nguyen, Naseem Khan, Khang Tran +2
The rapid rise of synthetic media has made deepfake detection a critical challenge for online safety and trust. Progress remains constrained by the scarcity of large, high-quality…
Program Structure-aware Language Models: Targeted Software Testing beyond Textual Semantics
Khang Tran, Khoa Nguyen, Cristian Borcea +1
Recent advances in large language models for test case generation have improved branch coverage via prompt-engineered mutations. However, they still lack principled mechanisms for…
Watermarking Degrades Alignment in Language Models: Analysis and Mitigation
Apurv Verma, NhatHai Phan, Shubhendu Trivedi
Watermarking has become a practical tool for tracing language model outputs, but it modifies token probabilities at inference time, which were carefully tuned by alignment training…