collaborators

8 papers

cs.LG2026

Multi-Objective Bayesian Optimization for Model Merging

Utkarsh Agarwal, Vamshi Bonagiri, Raul Astudillo +1

Model merging combines trained models directly in weight space, offering a compute-efficient alternative to additional fine-tuning. Selecting merge parameters is nevertheless diffi…

cs.CL2026

Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety

Vamshi Krishna Bonagiri, Ponnurangam Kumaragurum, Khanh Nguyen +1

As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical. While uncertainty quantification is w…

cs.AI2026

Cognitive Digital Twins: Ethical Risks and Governance for AI Systems That Model the Mind

Vamshi Krishna Bonagiri, Juan Nicolas Sepulveda-Arias, Abdoul Jalil Djiberou Mahamadou +1

As AI systems become increasingly persistent and personalized, they make possible a class of technologies that we call cognitive digital twins (CDTs): dynamic computational represe…

cs.CY2026

Efficient Safety Benchmarking via Item Response Theory

Fabio Spagliardi, Mírian Silva, Ayan Datta +3

Safety benchmarks for language models are typically evaluated using static paradigms that treat all items as equally informative for all models, an assumption that is particularly…

cs.CL2026

Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs

Krishak Aneja, Manas Mittal, Anmol Goel +2

Fine-tuning Large Language Models (LLMs) on benign narrow data can sometimes induce broad harmful behaviors, a vulnerability termed emergent misalignment (EM). While prior work lin…

cs.CL2026

Flying Pigs, FaR and Beyond: Evaluating LLM Reasoning in Counterfactual Worlds

Anish R Joishy, Ishwar B Balappanawar, Vamshi Krishna Bonagiri +3

A fundamental challenge in reasoning is navigating hypothetical, counterfactual worlds where logic may conflict with ingrained knowledge. We investigate this frontier for Large Lan…