activity
20232026
most citedThe Jumping Reasoning Curve? Tracking the Evolution of Reasoning Performance in GPT-[n] and o-[n] Models on Multimodal Puzzles

1 citations · 1 across the 9 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step

Vernon Toh, Navonil Majumder, Zhengyuan Liu +2

To operate robustly in open-world environments, autonomous agents should be able to infer the behavior of unfamiliar systems through interaction alone, even in the absence of docum…

cs.CL2026

GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards

Tej Deep Pala, Vernon Toh, Soujanya Poria

Reinforcement learning with verifiable rewards (e.g. GRPO) is now a common way to improve mathematical reasoning in Large Language Models (LLMs). However, current methods usually b…

cs.CL2025

Lessons from Training Grounded LLMs with Verifiable Rewards

Shang Hong Sim, Tej Deep Pala, Vernon Toh +5

Generating grounded and trustworthy responses remains a key challenge for large language models (LLMs). While retrieval-augmented generation (RAG) with citation-based grounding hol…

cs.CL2024

Not All Votes Count! Programs as Verifiers Improve Self-Consistency of Language Models for Math Reasoning

Vernon Y. H. Toh, Deepanway Ghosal, Soujanya Poria

Large language models (LLMs) have shown increasing competence in solving mathematical reasoning problems. However, many open-source LLMs still struggle with errors in calculation a…

cs.CL2024

Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique

Tej Deep Pala, Vernon Y. H. Toh, Rishabh Bhardwaj +1

In today's era, where large language models (LLMs) are integrated into numerous real-world applications, ensuring their safety and robustness is crucial for responsible AI usage. A…

cs.CL2023

VerityMath: Advancing Mathematical Reasoning by Self-Verification Through Unit Consistency

Vernon Toh Yan Han, Ratish Puduppully, Nancy F. Chen

Large Language Models (LLMs), combined with program-based solving techniques, are increasingly demonstrating proficiency in mathematical reasoning. For example, closed-source model…