16 papers
AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs
Hoang Nguyen, Sidharth Surapaneni, Akshay Kalkunte +9
Large Audio Language Models (LALMs) are rapidly advancing, but evaluating them remains challenging due to inefficient and non-standardized toolkits that limit fair comparison and s…
Super Apriel: One Checkpoint, Many Speeds
SLAM Labs, :, Oleksiy Ostapenko +13
We release Super Apriel, a 15B-parameter supernet in which every decoder layer provides four trained mixer choices -- Full Attention (FA), Sliding Window Attention (SWA), Kimi Delt…
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
Shiva Krishna Reddy Malay, Shravan Nayak, Jishnu Sethumadhavan Nair +6
Large language models are shifting from passive information providers to active agents intended for complex workflows. However, their deployment as reliable AI workers in enterpris…
AprielGuard
Jaykumar Kasundra, Anjaneya Praharaj, Sourabh Surana +11
Safeguarding large language models (LLMs) against unsafe or adversarial behavior is critical as they are increasingly deployed in conversational and agentic settings. Existing mode…
Grammar Search for Multi-Agent Systems
Mayank Singh, Vikas Yadav, Shiva Krishna Reddy Malay +4
Automatic search for Multi-Agent Systems has recently emerged as a key focus in agentic AI research. Several prior approaches have relied on LLM-based free-form search over the cod…
Apriel-H1: Towards Efficient Enterprise Reasoning Models
Oleksiy Ostapenko, Luke Kumar, Raymond Li +10
Large Language Models (LLMs) achieve remarkable reasoning capabilities through transformer architectures with attention mechanisms. However, transformers suffer from quadratic time…