9 citations · 20 across the 17 of their papers we have counts for
12 papers · 1 filter
ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning
Binglin Zhou, Peng Shi, Ryo Kamoi +2
Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tables, charts, and tex…
SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model
Xiaoxin Lu, Ranran Haoran Zhang, Rui Zhang
Large language models (LLMs) are increasingly deployed as planners for autonomous agents in household environments. While existing benchmarks evaluate whether LLM-generated plans e…
From Correctness to Preference: A Framework for Personalized Agentic Reinforcement Learning
Ranxu zhang, zeyang li, Jiacheng Huang +5
Agentic reinforcement learning (Agentic RL) has achieved strong progress in tasks with clear success signals. However, many real-world agent applications require user-conditioned b…
Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction
Ryo Kamoi, Ameya Godbole, Binglin Zhou +5
Simulating human conversations using large language models (LLMs) has emerged as a scalable methodology for modeling human social interaction. This paper reconsiders the evaluation…
Efficient PRM Training Data Synthesis via Formal Verification
Ryo Kamoi, Yusen Zhang, Nan Zhang +4
Process Reward Models (PRMs) have emerged as a promising approach for improving LLM reasoning capabilities by providing process supervision over reasoning traces. However, existing…
Verbosity Veracity: Demystify Verbosity Compensation Behavior of Large Language Models
Yusen Zhang, Sarkar Snigdha Sarathi Das, Rui Zhang
Although Large Language Models (LLMs) have demonstrated their strong capabilities in various tasks, recent work has revealed LLMs also exhibit undesirable behaviors, such as halluc…