activity
20242026
collaborators

6 papers

cs.CL2026

Fast-Slow Thinking RM: Efficient Integration of Scalar and Generative Reward Models

Jiayun Wu, Peixu Hou, Shan Qu +3

Reward models (RMs) are critical for aligning Large Language Models via Reinforcement Learning from Human Feedback (RLHF). While Generative Reward Models (GRMs) achieve superior ac…

cs.LG2025

Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-Correction

Yong Lin, Shange Tang, Bohan Lyu +17

We introduce Goedel-Prover-V2, a series of open-source language models that set a new state-of-the-art in automated theorem proving. Built on the standard expert iteration and rein…

cs.LG2025

Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving

Yong Lin, Shange Tang, Bohan Lyu +8

We introduce Goedel-Prover, an open-source language model that achieves state-of-the-art (as of April 5 2025) performance in automated formal proof generation for mathematical prob…

cs.HC2025

Adaptive Human-Agent Teaming: A Review of Empirical Studies from the Process Dynamics Perspective

Mengyao Wang, Jiayun Wu, Shuai Ma +4

The rapid advancement of AI, including Large Language Models, has propelled autonomous agents forward, accelerating the human-agent teaming (HAT) paradigm to leverage complementary…

cs.LG2024

Benign Overfitting in Out-of-Distribution Generalization of Linear Models

Shange Tang, Jiayun Wu, Jianqing Fan +1

Benign overfitting refers to the phenomenon where an over-parameterized model fits the training data perfectly, including noise in the data, but still generalizes well to the unsee…

cs.CL2024

Inductive-Deductive Strategy Reuse for Multi-Turn Instructional Dialogues

Jiao Ou, Jiayu Wu, Che Liu +3

Aligning large language models (LLMs) with human expectations requires high-quality instructional dialogues, which usually require instructions that are diverse and in-depth. Exist…