9 papers
ESF-Bench: Benchmarking Challenging Slot-Filling Scenarios for Real-World Enterprise Applications
Toby Liang, Gopal Sarda, Sagar Davasam +1
The rapid rise of large language models (LLMs) has driven transformative adoption across enterprises. However, deploying these models in real-world settings presents unique challen…
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding
Abhigya Verma, Khyati Mahajan, Amit Kumar Saha +4
Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MMLongBench-Doc. However, real-world documen…
R2V Agent: Teaching SLMs When to Ask for Help
Raghu Vamshi Hemadri, Humaira Firdowse Mohammed, Rishabh Maheshwary +5
Efficient agentic systems should incur expensive frontier-model costs only on decisions where a cheaper local model is likely to fail. Existing LLM cascades usually route whole que…
Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics
Jishnu Sethumadhavan Nair, Patrice Bechard, Rishabh Maheshwary +14
World models enable agents to anticipate the effects of their actions by internalizing environment dynamics. In enterprise systems, however, these dynamics are often defined by ten…
Terminal Agents Suffice for Enterprise Automation
Patrice Bechard, Orlando Marquez Ayala, Emily Chen +5
There has been growing interest in building agents that can interact with digital platforms to execute meaningful enterprise tasks autonomously. Among the approaches explored are t…
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
Shiva Krishna Reddy Malay, Shravan Nayak, Jishnu Sethumadhavan Nair +6
Large language models are shifting from passive information providers to active agents intended for complex workflows. However, their deployment as reliable AI workers in enterpris…