5 papers
Can Vision-Language Models Reason about AI Edits in Images?
Darsha Udayanga, Pin-Yu Chen, Payel Das +1
The paper explores training vision-language models with reinforcement learning to detect and localize AI-generated image edits, using reasoning traces and a lightweight segmentatio…
Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models
Huzaifa Arif, Keerthiram Murugesan, Ching-Yun Ko +3
We propose patching for large language models (LLMs) like software versions, a lightweight and modular approach for addressing safety vulnerabilities. While vendors release improve…
Combining Domain and Alignment Vectors to Achieve Better Knowledge-Safety Trade-offs in LLMs
Megh Thakkar, Quentin Fournier, Matthew Riemer +4
There is a growing interest in training domain-expert LLMs that excel in specific technical fields compared to their general-purpose instruction-tuned counterparts. However, these…
PEEL the Layers and Find Yourself: Revisiting Inference-time Data Leakage for Residual Neural Networks
Huzaifa Arif, Keerthiram Murugesan, Payel Das +2
This paper explores inference-time data leakage risks of deep neural networks (NNs), where a curious and honest model service provider is interested in retrieving users' private da…
Fundamental Safety-Capability Trade-offs in Fine-tuning Large Language Models
Pin-Yu Chen, Han Shen, Payel Das +1
Fine-tuning Large Language Models (LLMs) on some task-specific datasets has been a primary use of LLMs. However, it has been empirically observed that this approach to enhancing ca…