activity
20242026
collaborators

7 papers

cs.AI2026

Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI Applications

Nathalie Baracaldo, Nicolas Mello, Kush R. Varshney +3

When it comes to safety policies for generative AI, one size does not fit all. Each organization and use case needs to mitigate different risks depending on the application context…

cs.LG2026

Balancing Multi-modal Sensor Learning via Multi-objective Optimization

Heshan Fernando, Quan Xiao, Parikshit Ram +4

Learning-enabled control systems increasingly rely on multiple sensing modalities (e.g., vision, audio, language, etc.) for perception and decision support. A key challenge is that…

cs.LG2025

Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective

Heshan Fernando, Han Shen, Parikshit Ram +4

The post-training of LLMs, which typically consists of the supervised fine-tuning (SFT) stage and the preference learning stage (RLHF or DPO), is crucial to effective and safe LLM…

cs.CR2025

Towards a Re-evaluation of Data Forging Attacks in Practice

Mohamed Suliman, Anisa Halimi, Swanand Kadhe +2

Data forging attacks provide counterfactual proof that a model was trained on a given dataset, when in fact, it was trained on another. These attacks work by forging (replacing) mi…

cs.AI2024

MAP: Multi-Human-Value Alignment Palette

Xinran Wang, Qi Le, Ammar Ahmed +5

Ensuring that generative AI systems align with human values is essential but challenging, especially when considering multiple human values and their potential trade-offs. Since hu…

cs.CR2024

Turning Generative Models Degenerate: The Power of Data Poisoning Attacks

Shuli Jiang, Swanand Ravindra Kadhe, Yi Zhou +3

The increasing use of large language models (LLMs) trained by third parties raises significant security concerns. In particular, malicious actors can introduce backdoors through po…