16 papers
Paved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training Regimes
Jeremias Ferrao, Niclas Müller-Hof, Iustin Sîrbu +2
We argue that safety classifiers should model user intent as an explicit signal between the prompt and the final label. To study this, we introduce AIMS, a human-annotated dataset…
Defenses & Enablers For Skill Injection Attacks on Terminal Based Agents
Yoshinari Fujinuma, Varun Gangal, Traian Rebedea +4
Large language model (LLM) agents increasingly rely on reusable skills i.e. documents describing task-specific procedures. However, this introduces a new attack surface for agents…
"ÃnÅ£elegi RomâneÅte?'' A Recipe for Romanian Vision-Language Models
Mihai Masala, Marius Leordeanu, Mihai Dascalu +1
Vision-Language Models (VLMs) largely follow the text-only LLM trajectory, excelling on English benchmarks but sharply degrading on low-resource languages, where neither large-scal…
SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair
Ion George Dinu, Marian Cristian MihÄescu, Traian Rebedea
Architectural code smells erode software maintainability and are costly to repair manually, yet unlike localized bugs, they require cross-module reasoning about design intent that…
Training a General Purpose Automated Red Teaming Model
Aishwarya Padmakumar, Leon Derczynski, Traian Rebedea +1
Automated methods for red teaming LLMs are an important tool to identify LLM vulnerabilities that may not be covered in static benchmarks, allowing for more thorough probing. They…
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
NVIDIA, :, Aakshita Chandiramani +544
We describe the pre-training, post-training, and quantization of Nemotron 3 Super, a 120 billion (active 12 billion) parameter hybrid Mamba-Attention Mixture-of-Experts model. Nemo…