Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF
Yuan Fang, Yiming Luo, Aimin Zhou +1
Ensuring the safety of large language models (LLMs) requires robust red teaming, yet the systematic synthesis of high-quality toxic data remains under-explored. We propose Reverse…
cs.CL2025
Consultant Decoding: Yet Another Synergistic Mechanism
Chuanghao Ding, Jiaping Wang, Ziqing Yang +4
The synergistic mechanism based on Speculative Decoding (SD) has garnered considerable attention as a simple yet effective approach for accelerating the inference of large language…
cs.CL2024
Reward Difference Optimization For Sample Reweighting In Offline RLHF
Shiqi Wang, Zhengze Zhang, Rui Zhao +2
With the rapid advances in Large Language Models (LLMs), aligning LLMs with human preferences become increasingly important. Although Reinforcement Learning with Human Feedback (RL…