3 citations · 3 across the 13 of their papers we have counts for
6 papers · 1 filter
Tandem Reinforcement Learning with Verifiable Rewards
Difan Jiao, Raghav Singhal, Robert West +1
Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capability of large language models, reaching expert or even superhuman performance i…
LLM Safety From Within: Detecting Harmful Content with Internal Representations
Difan Jiao, Yilun Liu, Ye Yuan +4
Guard models are widely used to detect harmful content in user prompts and LLM responses. However, state-of-the-art guard models rely solely on terminal-layer representations and o…
ThinkTwice: Jointly Optimizing Large Language Models for Reasoning and Self-Refinement
Difan Jiao, Qianfeng Wen, Blair Yang +2
We introduce ThinkTwice, a simple two-phase framework that jointly optimizes LLMs to solve reasoning problems and refine the answers, based on Group Relative Policy Optimization (G…
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
Zhenwei Tang, Difan Jiao, Blair Yang +1
Evaluating whether vision-language models (VLMs) reason consistently across representations is challenging because modality comparisons are typically confounded by task differences…
Learning to Imitate with Less: Efficient Individual Behavior Modeling in Chess
Zhenwei Tang, Difan Jiao, Eric Xue +4
As humans seek to collaborate with, learn from, and better understand artificial intelligence systems, developing AIs that can accurately emulate individual decision-making becomes…
Maia-2: A Unified Model for Human-AI Alignment in Chess
Zhenwei Tang, Difan Jiao, Reid McIlroy-Young +3
There are an increasing number of domains in which artificial intelligence (AI) systems both surpass human ability and accurately model human behavior. This introduces the possibil…