collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

SEIF: Self-Evolving Reinforcement Learning for Instruction Following

Qingyu Ren, Qianyu He, Jiajie Zhu +7

Instruction following is a fundamental capability of large language models (LLMs), yet continuously improving this capability remains challenging. Existing methods typically rely e…

cs.CL2024

Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data

Han Xia, Songyang Gao, Qiming Ge +3

Reinforcement Learning from Human Feedback (RLHF) has proven effective in aligning large language models with human intentions, yet it often relies on complex methodologies like Pr…

cs.CL20246 cited

EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models

Weikang Zhou, Xiao Wang, Limao Xiong +18

Jailbreak attacks are crucial for identifying and mitigating the security vulnerabilities of Large Language Models (LLMs). They are designed to bypass safeguards and elicit prohibi…

cs.CL2024

RoCoIns: Enhancing Robustness of Large Language Models through Code-Style Instructions

Yuansen Zhang, Xiao Wang, Zhiheng Xi +4

Large Language Models (LLMs) have showcased remarkable capabilities in following human instructions. However, recent studies have raised concerns about the robustness of LLMs when…

cs.CL20231 cited

Orthogonal Subspace Learning for Language Model Continual Learning

Xiao Wang, Tianze Chen, Qiming Ge +6

Benefiting from massive corpora and advanced hardware, large language models (LLMs) exhibit remarkable capabilities in language understanding and generation. However, their perform…