9 papers
ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step
Vernon Toh, Navonil Majumder, Zhengyuan Liu +2
To operate robustly in open-world environments, autonomous agents should be able to infer the behavior of unfamiliar systems through interaction alone, even in the absence of docum…
GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewards
Tej Deep Pala, Vernon Toh, Soujanya Poria
Reinforcement learning with verifiable rewards (e.g. GRPO) is now a common way to improve mathematical reasoning in Large Language Models (LLMs). However, current methods usually b…
Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned
Brandon Ong, Tej Deep Pala, Vernon Toh +2
Process Reward Models (PRMs) provide step-level supervision that improves the reliability of reasoning in large language models. While PRMs have been extensively studied in text-ba…
Lessons from Training Grounded LLMs with Verifiable Rewards
Shang Hong Sim, Tej Deep Pala, Vernon Toh +5
Generating grounded and trustworthy responses remains a key challenge for large language models (LLMs). While retrieval-augmented generation (RAG) with citation-based grounding hol…
The Jumping Reasoning Curve? Tracking the Evolution of Reasoning Performance in GPT-[n] and o-[n] Models on Multimodal Puzzles
Vernon Y. H. Toh, Yew Ken Chia, Deepanway Ghosal +1
The releases of OpenAI's o-[n] series, such as o1, o3, and o4-mini, mark a significant paradigm shift in Large Language Models towards advanced reasoning capabilities. Notably, mod…
Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
Qi Sun, Pengfei Hong, Tej Deep Pala +4
Traditional reinforcement learning-based robotic control methods are often task-specific and fail to generalize across diverse environments or unseen objects and instructions. Visu…