activity
20232026
most citedChain of Agents: Large Language Models Collaborating on Long-Context Tasks

9 citations · 20 across the 17 of their papers we have counts for

collaborators
Showing cs.CLShow all

12 papers · 1 filter

cs.CL2026

ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning

Binglin Zhou, Peng Shi, Ryo Kamoi +2

Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tables, charts, and tex…

cs.CL2026

SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model

Xiaoxin Lu, Ranran Haoran Zhang, Rui Zhang

Large language models (LLMs) are increasingly deployed as planners for autonomous agents in household environments. While existing benchmarks evaluate whether LLM-generated plans e…

cs.CL2026

From Correctness to Preference: A Framework for Personalized Agentic Reinforcement Learning

Ranxu zhang, zeyang li, Jiacheng Huang +5

Agentic reinforcement learning (Agentic RL) has achieved strong progress in tasks with clear success signals. However, many real-world agent applications require user-conditioned b…

cs.CL2026

Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction

Ryo Kamoi, Ameya Godbole, Binglin Zhou +5

Simulating human conversations using large language models (LLMs) has emerged as a scalable methodology for modeling human social interaction. This paper reconsiders the evaluation…

cs.CL2025

Efficient PRM Training Data Synthesis via Formal Verification

Ryo Kamoi, Yusen Zhang, Nan Zhang +4

Process Reward Models (PRMs) have emerged as a promising approach for improving LLM reasoning capabilities by providing process supervision over reasoning traces. However, existing…

cs.CL20242 cited

Verbosity Veracity: Demystify Verbosity Compensation Behavior of Large Language Models

Yusen Zhang, Sarkar Snigdha Sarathi Das, Rui Zhang

Although Large Language Models (LLMs) have demonstrated their strong capabilities in various tasks, recent work has revealed LLMs also exhibit undesirable behaviors, such as halluc…