3 papers
cs.LG2026
Evaluating Robustness of Large Language Models in Enterprise Applications: Benchmarks for Perturbation Consistency Across Formats and Languages
Tara Bogavelli, Oluwanifemi Bamgbose, Gabrielle Gauthier Melançon +2
Enterprise LLM applications require consistently high quality and reliable performance across diverse scenarios, demanding robustness to minor variations. Existing research shows t…
cs.AI2026
AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
Tara Bogavelli, Roshnee Sharma, Hari Subramani
While individual components of agentic architectures have been studied in isolation, there remains limited empirical understanding of how different design dimensions interact withi…
cs.LG2025
Revitalizing Saturated Benchmarks: A Weighted Metric Approach for Differentiating Large Language Model Performance
Bryan Etzine, Masoud Hashemi, Nishanth Madhusudhan +4
Existing benchmarks are becoming saturated and struggle to separate model performances due to factors like data contamination and advancing LLM capabilities. This paper introduces…