5 papers
Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions
Oluwanifemi Bamgbose, Simon Rosen, Jash Shah +6
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception,…
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
Tara Bogavelli, Gabrielle Gauthier Melançon, Katrina Stankiewicz +10
Voice agents, artificial intelligence systems that conduct spoken conversations to complete tasks, are increasingly deployed across enterprise applications. However, no existing be…
Evaluating Robustness of Large Language Models in Enterprise Applications: Benchmarks for Perturbation Consistency Across Formats and Languages
Tara Bogavelli, Oluwanifemi Bamgbose, Gabrielle Gauthier Melançon +2
Enterprise LLM applications require consistently high quality and reliable performance across diverse scenarios, demanding robustness to minor variations. Existing research shows t…
AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
Tara Bogavelli, Roshnee Sharma, Hari Subramani
While individual components of agentic architectures have been studied in isolation, there remains limited empirical understanding of how different design dimensions interact withi…
Apriel-Nemotron-15B-Thinker
Shruthan Radhakrishna, Soham Parikh, Gopal Sarda +32
While large language models (LLMs) have achieved remarkable reasoning capabilities across domains like code, math and other enterprise tasks, their significant memory and computati…