4 papers
MM-tau-p: Persona-Adaptive Prompting for Robust Multi-Modal Agent Evaluation in Dual-Control Settings
Anupam Purwar, Aditya Choudhary
Current evaluation frameworks and benchmarks for LLM powered agents focus on text chat driven agents, these frameworks do not expose the persona of user to the agent, thus operatin…
When Fine-Tuning Fails and when it Generalises: Role of Data Diversity and Mixed Training in LLM-based TTS
Anupam Purwar, Aditya Choudhary
Large language models are increasingly adopted as semantic backbones for neural text-to-speech systems. However, frozen LLM representations are insufficient for modeling speaker sp…
FOCAL: A Novel Benchmarking Technique for Multi-modal Agents
Anupam Purwar, Aditya Choudhary
With the recent advancements in reasoning capabilities, tool calling using MCP servers and Audio Language Models (ALMs), development and integration of multi-modal agents (with voi…
i-LAVA: Insights on Low Latency Voice-2-Voice Architecture for Agents
Anupam Purwar, Aditya Choudhary
We experiment with a low-latency, end-to-end voice-to-voice communication model to optimize it for real-time conversational applications. By analyzing components essential to voice…