2 papers
cs.CL2025
NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
NVIDIA, :, Aarti Basant +214
We introduce Nemotron-Nano-9B-v2, a hybrid Mamba-Transformer language model designed to increase throughput for reasoning workloads while achieving state-of-the-art accuracy compar…
cs.LG2025
SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs
Shaona Ghosh, Amrita Bhattacharjee, Yftah Ziser +1
Fine-tuning large language models (LLMs) to adapt to evolving safety policies is costly and impractical. Mechanistic interpretability enables inference-time control through latent…