7 papers
Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI Applications
Nathalie Baracaldo, Nicolas Mello, Kush R. Varshney +3
When it comes to safety policies for generative AI, one size does not fit all. Each organization and use case needs to mitigate different risks depending on the application context…
Balancing Multi-modal Sensor Learning via Multi-objective Optimization
Heshan Fernando, Quan Xiao, Parikshit Ram +4
Learning-enabled control systems increasingly rely on multiple sensing modalities (e.g., vision, audio, language, etc.) for perception and decision support. A key challenge is that…
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
Heshan Fernando, Han Shen, Parikshit Ram +4
The post-training of LLMs, which typically consists of the supervised fine-tuning (SFT) stage and the preference learning stage (RLHF or DPO), is crucial to effective and safe LLM…
Towards a Re-evaluation of Data Forging Attacks in Practice
Mohamed Suliman, Anisa Halimi, Swanand Kadhe +2
Data forging attacks provide counterfactual proof that a model was trained on a given dataset, when in fact, it was trained on another. These attacks work by forging (replacing) mi…
MAP: Multi-Human-Value Alignment Palette
Xinran Wang, Qi Le, Ammar Ahmed +5
Ensuring that generative AI systems align with human values is essential but challenging, especially when considering multiple human values and their potential trade-offs. Since hu…
Turning Generative Models Degenerate: The Power of Data Poisoning Attacks
Shuli Jiang, Swanand Ravindra Kadhe, Yi Zhou +3
The increasing use of large language models (LLMs) trained by third parties raises significant security concerns. In particular, malicious actors can introduce backdoors through po…