collaborators

5 papers

cs.LG2026

Deliberative Alignment is Deep, but Uncertainty Remains: Inference time safety improvement in reasoning via attribution of unsafe behavior to base model

Pankayaraj Pathmanathan, Furong Huang

While the wide adoption of refusal training in large language models (LLMs) has showcased improvements in model safety, recent works have highlighted shortcomings due to the shallo…

cs.AI2026

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?

Udari Madhushani Sehwag, Elaine Lau, Haniyeh Ehsani Oskouie +14

Accelerating scientific discovery requires the identification of which experiments would yield the best outcomes before committing resources to costly physical validation. While ex…

cs.AI2026

Agentic Critical Training

Weize Liu, Minghui Liu, Sy-Tuyen Ho +3

Training large language models (LLMs) as autonomous agents often begins with imitation learning, but it only teaches agents what to do without understanding why: agents never contr…

cs.CL2025

Hold Onto That Thought: Assessing KV Cache Compression On Reasoning

Minghui Liu, Aadi Palnitkar, Tahseen Rabbani +9

Large language models (LLMs) have demonstrated remarkable performance on long-context tasks, but are often bottlenecked by memory constraints. Namely, the KV cache, which is used t…

cs.CY2025

PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach

Udari Madhushani Sehwag, Shayan Shabihi, Alex McAvoy +4

Recent advances in Large Language Models (LLMs) have sparked concerns over their potential to acquire and misuse dangerous or high-risk capabilities, posing frontier risks. Current…