3 papers
cs.LG2026
Evaluating Robustness of Large Language Models in Enterprise Applications: Benchmarks for Perturbation Consistency Across Formats and Languages
Tara Bogavelli, Oluwanifemi Bamgbose, Gabrielle Gauthier Melançon +2
Enterprise LLM applications require consistently high quality and reliable performance across diverse scenarios, demanding robustness to minor variations. Existing research shows t…
cs.AI2025
AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
Tara Bogavelli, Roshnee Sharma, Hari Subramani
While individual components of agentic architectures have been studied in isolation, there remains limited empirical understanding of how different design dimensions interact withi…
cs.LG2025
Apriel-Nemotron-15B-Thinker
Shruthan Radhakrishna, Soham Parikh, Gopal Sarda +32
While large language models (LLMs) have achieved remarkable reasoning capabilities across domains like code, math and other enterprise tasks, their significant memory and computati…