14 papers
GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings
Shivali Dalmia, Sumukha Thoppanahalli, Mohammadreza Sediqin +1
Enterprise guideline documents are heterogeneous and multimodal, combining narrative text, complex tables, and embedded images. Existing LLM and VLM systems face hallucinated conte…
DuplexWorld: Can voice agents help you get through the day?
Aryan Vijay Bhosale, Harshit Rajgarhia, Akhil Pothanapalli +3
Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversatio…
TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models
Aryan Vijay Bhosale, Harshit Rajgarhia, Abhishek Mukherji +1
Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a basic question about them remains unanswered…
CaReCoS: A Spectrogram based Visual Benchmark for Cardiac, Respiratory and Cough Sounds
Harshit Rajgarhia, Shuubham Ojha, Akhil Pothanapalli +4
Medical acoustic signals such as respiratory sounds, cardiac auscultations, and cough audio carry rich diagnostic information, yet no existing benchmark evaluates multimodal reason…
World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments
Ananya Mantravadi, Harshit Rajgarhia, Prasanna Desikan +1
Clinical protocol-execution tasks -- checking a lab value, applying a threshold, placing a correctly structured FHIR order -- are natural candidates for RL from world feedback: onc…
Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows
Karthikeya Aditya Vissa, Sankalp Mane, Ananya Mantravadi +2
Large language models are trained to predict the next token, not to act inside a specific API. In niche enterprise SaaS workflows -- where success means hitting the right endpoint…