activity
20242026
most citedEnhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization

2 citations · 2 across the 7 of their papers we have counts for

collaborators

7 papers

cs.LG2026

A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics

Zheshun Wu, Renjie Zheng, Jinhang Zuo +2

This paper investigates a hybrid reinforcement learning setting in tabular Markov Decision Processes (MDPs), where an agent aims to learn an optimal policy by combining online inte…

cs.LG2026

Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents

Jiawei Huang, Qingping Yang, Renjie Zheng +1

Despite recent progress in Large Language Model (LLM) Agents for Software Engineering (SWE) tasks, end-to-end fine-tuning typically relies on verifiable terminal rewards such as wh…

cs.AI2025

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Haoming Wang, Haoyang Zou, Huatong Song +109

The development of autonomous agents for graphical user interfaces (GUIs) presents major challenges in artificial intelligence. While recent advances in native agent models have sh…

cs.CL2025

SciDA: Scientific Dynamic Assessor of LLMs

Junting Zhou, Tingjia Miao, Yiyan Liao +15

Advancement in Large Language Models (LLMs) reasoning capabilities enables them to solve scientific problems with enhanced efficacy. Thereby, a high-quality benchmark for comprehen…

cs.LG2024

Flaming-hot Initiation with Regular Execution Sampling for Large Language Models

Weizhe Chen, Zhicheng Zhang, Guanlin Liu +6

Since the release of ChatGPT, large language models (LLMs) have demonstrated remarkable capabilities across various domains. A key challenge in developing these general capabilitie…

cs.AI2024

Process Supervision-Guided Policy Optimization for Code Generation

Ning Dai, Zheng Wu, Renjie Zheng +7

Reinforcement learning (RL) with unit test feedback has enhanced large language models' (LLMs) code generation, but relies on sparse rewards provided only after complete code evalu…