most citedTraining Language Models to Self-Correct via Reinforcement Learning

7 citations · 7 across the 3 of their papers we have counts for

collaborators

6 papers

cs.SE2025

SmartMLOps Studio: Design of an LLM-Integrated IDE with Automated MLOps Pipelines for Model Development and Monitoring

Jiawei Jin, Yingxin Su, Xiaotong Zhu

The rapid expansion of artificial intelligence and machine learning (ML) applications has intensified the demand for integrated environments that unify model development, deploymen…

cs.DC2025

Design and Implementation of Code Completion System Based on LLM and CodeBERT Hybrid Subsystem

Bingbing Zhang, Ziyu Lin, Yingxin Su

In the rapidly evolving industry of software development, coding efficiency and accuracy play significant roles in delivering high-quality software. Various code suggestion and com…

cs.CE2025

Aethorix v1.0: An Integrated Scientific AI Agent for Scalable Inorganic Materials Innovation and Industrial Implementation

Yingjie Shi, Yiru Gong, Yiqun Su +3

Artificial Intelligence (AI) is redefining the frontiers of scientific domains, ranging from drug discovery to meteorological modeling, yet its integration within industrial manufa…

cs.CR2025

Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

Yakai Li, Jiekang Hu, Weiduan Sang +7

Large Language Models face security threats from jailbreak attacks. Existing research has predominantly focused on prompt-level attacks while largely ignoring the underexplored att…

eess.AS2024

Device-Directed Speech Detection for Follow-up Conversations Using Large Language Models

Ognjen, Rudovic, Pranay Dighe +7

Follow-up conversations with virtual assistants (VAs) enable a user to seamlessly interact with a VA without the need to repeatedly invoke it using a keyword (after the first query…

cs.LG20247 cited

Training Language Models to Self-Correct via Reinforcement Learning

Aviral Kumar, Vincent Zhuang, Rishabh Agarwal +15

Self-correction is a highly desirable capability of large language models (LLMs), yet it has consistently been found to be largely ineffective in modern LLMs. Current methods for t…