10 papers
Hey Chat, Can You Teach Me? Structuring Socratic Dialogue for Human Learning in the Wild
Sidney Tio, Arunesh Sinha, Pradeep Varakantham
Large language models are now widely used for everyday learning, but the underlying interactions are typically unstructured chats rather than following a curriculum. Unlike formal…
Robust Critics: Defending LLMs Against Multi-Turn Attacks
Roman Belaire, Arunesh Sinha, Pradeep Varakantham
When a user asks a language model something harmful, is it a genuine attack or a misunderstood but well-meaning question? This ambiguity is one of the central challenges of LLM saf…
Induced Numerical Instability: Hidden Costs in Multimodal Large Language Models
Wai Tuck Wong, Jun Sun, Arunesh Sinha
The use of multimodal large language models has become widespread, and as such the study of these models and their failure points has become of utmost importance. We study a novel…
Strategic Incentivization for Locally Differentially Private Federated Learning
Yashwant Krishna Pagoti, Arunesh Sinha, Shamik Sural
In Federated Learning (FL), multiple clients jointly train a machine learning model by sharing gradient information, instead of raw data, with a server over multiple rounds. To add…
Automatic LLM Red Teaming
Roman Belaire, Arunesh Sinha, Pradeep Varakantham
Red teaming is critical for identifying vulnerabilities and building trust in current LLMs. However, current automated methods for Large Language Models (LLMs) rely on brittle prom…
On Minimizing Adversarial Counterfactual Error in Adversarial RL
Roman Belaire, Arunesh Sinha, Pradeep Varakantham
Deep Reinforcement Learning (DRL) policies are highly susceptible to adversarial noise in observations, which poses significant risks in safety-critical scenarios. The challenge in…