1 citations · 1 across the 9 of their papers we have counts for
6 papers · 1 filter
ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step
Vernon Toh, Navonil Majumder, Zhengyuan Liu +2
To operate robustly in open-world environments, autonomous agents should be able to infer the behavior of unfamiliar systems through interaction alone, even in the absence of docum…
GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards
Tej Deep Pala, Vernon Toh, Soujanya Poria
Reinforcement learning with verifiable rewards (e.g. GRPO) is now a common way to improve mathematical reasoning in Large Language Models (LLMs). However, current methods usually b…
Lessons from Training Grounded LLMs with Verifiable Rewards
Shang Hong Sim, Tej Deep Pala, Vernon Toh +5
Generating grounded and trustworthy responses remains a key challenge for large language models (LLMs). While retrieval-augmented generation (RAG) with citation-based grounding hol…
Not All Votes Count! Programs as Verifiers Improve Self-Consistency of Language Models for Math Reasoning
Vernon Y. H. Toh, Deepanway Ghosal, Soujanya Poria
Large language models (LLMs) have shown increasing competence in solving mathematical reasoning problems. However, many open-source LLMs still struggle with errors in calculation a…
Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique
Tej Deep Pala, Vernon Y. H. Toh, Rishabh Bhardwaj +1
In today's era, where large language models (LLMs) are integrated into numerous real-world applications, ensuring their safety and robustness is crucial for responsible AI usage. A…
VerityMath: Advancing Mathematical Reasoning by Self-Verification Through Unit Consistency
Vernon Toh Yan Han, Ratish Puduppully, Nancy F. Chen
Large Language Models (LLMs), combined with program-based solving techniques, are increasingly demonstrating proficiency in mathematical reasoning. For example, closed-source model…