4 papers · 1 filter
Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models
Raj Jaiswal, Dhruv Jain, Rishabh Dhawan +4
Physics reasoning fails structurally in small language models: an error at any step propagates forward, corrupting every inference that follows. Limited domain knowledge, hallucina…
VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
Dhruv Jain, Harshit Shukla, Gautam Rajeev +3
Large scale Speech Language Models have enabled voice assistants capable of understanding natural spoken queries and performing complex tasks. However, existing speech benchmarks l…
Improving Physics Reasoning in Large Language Models Using Mixture of Refinement Agents
Raj Jaiswal, Dhruv Jain, Harsh Parimal Popat +4
Large Language Models (LLMs) demonstrate remarkable capabilities in various reasoning tasks. However, they encounter significant challenges when it comes to scientific reasoning, p…
SwissNYF: Tool Grounded LLM Agents for Black Box Setting
Somnath Sendhil Kumar, Dhruv Jain, Eshaan Agarwal +1
While Large Language Models (LLMs) have demonstrated enhanced capabilities in function-calling, these advancements primarily rely on accessing the functions' responses. This method…