6 papers
Forethought: Verifiable Reasoning from Neurosymbolic Primitive Programming
Vishvesh Bhat, Jay Vaghasiya, Emmanuel Anaya Gonzalez
Current agentic workflows usually involve decomposing user requests into sequences of tool calls with correctly resolved parameters, the results of which are processed through reas…
Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation
Vishvesh Bhat, Jay Vaghasiya, Muhammad Ahmed Mohsin +1
Tool-calling benchmarks are increasingly used to rank language-model agents, yet their scores are often treated as ground truth without validating the evaluators themselves. We pre…
MAVEN: Improving Generalization in Agentic Tool Calling
Omkar Ghugarkar, Vishvesh Bhat, Muhammad Ahmed Mohsin +1
Generalization across agentic tool-calling environments remains a central challenge for reliable agentic reasoning systems. Although large language models achieve strong results on…
Compositional Neuro-Symbolic Reasoning
Anugyan Das, Omkar Ghugarkar, Vishvesh Bhat +1
We study structured abstraction-based reasoning for the Abstraction and Reasoning Corpus (ARC) and compare its generalization to test-time approaches. Purely neural architectures l…
On Generalization in Agentic Tool Calling: CoreThink Agentic Reasoner and MAVEN Dataset
Vishvesh Bhat, Omkar Ghugarkar, Julian McAuley
Generalization across Agentic tool-calling environments remains a key unsolved challenge in developing reliable agentic reasoning systems. While large language models (LLMs) demons…
CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs
Jay Vaghasiya, Omkar Ghugarkar, Vishvesh Bhat +2
We introduce CoreThink, a state-of-the-art Reasoning Layer built upon a novel reasoning method called General Symbolics. This approach diverges from reasoning paradigms such as tes…