activity
20242026
collaborators

9 papers

cs.CY2026

Next-Billion AI Index: The compass for AI utility and adoption in the global majority

Ambrish Rawat, Jessica He, Subhabrata Majumdar +6

Generative AI assessments remain dominated by frontier capability benchmarks that often fail to capture whether systems can be sustainably deployed, adapted, and trusted in locally…

cs.LG2025

Building a Foundational Guardrail for General Agentic Systems via Synthetic Data

Yue Huang, Hang Hua, Yujun Zhou +11

While LLM agents can plan multi-step tasks, intervening at the planning stage-before any action is executed-is often the safest way to prevent harm, since certain risks can lead to…

cs.LG2025

Activated LoRA: Fine-tuned LLMs for Intrinsics

Kristjan Greenewald, Luis Lastras, Thomas Parnell +6

Low-Rank Adaptation (LoRA) has emerged as a highly efficient framework for finetuning the weights of large foundation models, and has become the go-to method for data-driven custom…

cs.CY2025

AI Risk Atlas: Taxonomy and Tooling for Navigating AI Risks and Resources

Frank Bagehorn, Kristina Brimijoin, Elizabeth M. Daly +17

The rapid evolution of generative AI has expanded the breadth of risks associated with AI systems. While various taxonomies and frameworks exist to classify these risks, the lack o…

cs.LG2025

MAD-MAX: Modular And Diverse Malicious Attack MiXtures for Automated LLM Red Teaming

Stefan Schoepf, Muhammad Zaid Hameed, Ambrish Rawat +4

With LLM usage rapidly increasing, their vulnerability to jailbreaks that create harmful outputs are a major security risk. As new jailbreaking strategies emerge and models are cha…

cs.CR2025

Attention Tracker: Detecting Prompt Injection Attacks in LLMs

Kuo-Han Hung, Ching-Yun Ko, Ambrish Rawat +3

Large Language Models (LLMs) have revolutionized various domains but remain vulnerable to prompt injection attacks, where malicious inputs manipulate the model into ignoring origin…