4 papers
Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment
Kartik Pandit, Sourav Ganguly, Arnesh Banerjee +2
Ensuring safety is a foundational requirement for large language models (LLMs). Achieving an appropriate balance between enhancing the utility of model outputs and mitigating their…
Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees
Sourav Ganguly, Kartik Pandit, Arnob Ghosh
Real-world decision-making systems operate in environments where state transitions depend not only on the agent's actions, but also on \textbf{exogenous factors outside its control…
Medical Image De-Identification Benchmark Challenge
Linmin Pei, Granger Sutton, Michael Rutherford +67
The de-identification (deID) of protected health information (PHI) and personally identifiable information (PII) is a fundamental requirement for sharing medical images, particular…
TPU-Gen: LLM-Driven Custom Tensor Processing Unit Generator
Deepak Vungarala, Mohammed E. Elbtity, Sumiya Syed +5
The increasing complexity and scale of Deep Neural Networks (DNNs) necessitate specialized tensor accelerators, such as Tensor Processing Units (TPUs), to meet various computationa…