activity
20172026
most citedA Survey on Neural Architecture Search

208 citations · 389 across the 22 of their papers we have counts for

collaborators
Showing cs.LGShow all

14 papers · 1 filter

cs.LG2025

Building a Foundational Guardrail for General Agentic Systems via Synthetic Data

Yue Huang, Hang Hua, Yujun Zhou +11

While LLM agents can plan multi-step tasks, intervening at the planning stage-before any action is executed-is often the safest way to prevent harm, since certain risks can lead to…

cs.LG2025

Activated LoRA: Fine-tuned LLMs for Intrinsics

Kristjan Greenewald, Luis Lastras, Thomas Parnell +6

Low-Rank Adaptation (LoRA) has emerged as a highly efficient framework for finetuning the weights of large foundation models, and has become the go-to method for data-driven custom…

cs.LG2025

MAD-MAX: Modular And Diverse Malicious Attack MiXtures for Automated LLM Red Teaming

Stefan Schoepf, Muhammad Zaid Hameed, Ambrish Rawat +4

With LLM usage rapidly increasing, their vulnerability to jailbreaks that create harmful outputs are a major security risk. As new jailbreaking strategies emerge and models are cha…

cs.LG2024

Detectors for Safe and Reliable LLMs: Implementations, Uses, and Limitations

Swapnaja Achintalwar, Adriana Alvarado Garcia, Ateret Anaby-Tavor +35

Large language models (LLMs) are susceptible to a variety of risks, from non-faithful output to biased and toxic generations. Due to several limiting factors surrounding LLMs (trai…

cs.LG20241 cited

Domain Adaptation for Time series Transformers using One-step fine-tuning

Subina Khanal, Seshu Tirupathi, Giulio Zizzo +2

The recent breakthrough of Transformers in deep learning has drawn significant attention of the time series community due to their ability to capture long-range dependencies. Howev…

cs.LG2023

FairSISA: Ensemble Post-Processing to Improve Fairness of Unlearning in LLMs

Swanand Ravindra Kadhe, Anisa Halimi, Ambrish Rawat +1

Training large language models (LLMs) is a costly endeavour in terms of time and computational resources. The large amount of training data used during the unsupervised pre-trainin…