activity
20242026
most citedTDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference

1 citations · 1 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CL2026

Talk Less, Verify More: Improving LLM Assistants with Semantic Checks and Execution Feedback

Yan Sun, Ming Cai, Stanley Kok

As large language model (LLM) assistants become increasingly integrated into enterprise workflows, their ability to generate accurate, semantically aligned, and executable outputs…

cs.LG20251 cited

TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference

Dan Zhang, Min Cai, Jonathan Light +3

Reward models are central to both reinforcement learning (RL) with language models and inference-time verification. However, existing reward models often lack temporal consistency,…

cs.LG2025

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment

Jiazheng Zhang, Wenqing Jing, Zizhuo Zhang +9

Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human values. However, noisy preferences in human feedback can lead to reward misgeneralizatio…

cs.CL2025

How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence

Hongzhe Du, Weikai Li, Min Cai +5

Post-training is essential for the success of large language models (LLMs), transforming pre-trained base models into more useful and aligned post-trained models. While plenty of w…

cs.CL2025

DataSciBench: An LLM Agent Benchmark for Data Science

Dan Zhang, Sining Zhoubian, Min Cai +7

This paper presents DataSciBench, a comprehensive benchmark for evaluating Large Language Model (LLM) capabilities in data science. Recent related benchmarks have primarily focused…

cs.AI2024

PIANIST: Learning Partially Observable World Models with LLMs for Multi-Agent Decision Making

Jonathan Light, Sixue Xing, Yuanzhe Liu +7

Effective extraction of the world knowledge in LLMs for complex decision-making tasks remains a challenge. We propose a framework PIANIST for decomposing the world model into seven…