Showing cs.CLShow all
3 papers · 1 filter
cs.CL2024
It Takes Two: On the Seamlessness between Reward and Policy Model in RLHF
Taiming Lu, Lingfeng Shen, Xinyu Yang +3
Reinforcement Learning from Human Feedback (RLHF) involves training policy models (PMs) and reward models (RMs) to align language models with human preferences. Instead of focusing…
cs.CL2024
DiffNorm: Self-Supervised Normalization for Non-autoregressive Speech-to-speech Translation
Weiting Tan, Jingyu Zhang, Lingfeng Shen +2
Non-autoregressive Transformers (NATs) are recently applied in direct speech-to-speech translation systems, which convert speech across different languages without intermediate tex…
cs.CL2023
SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation
Abe Bohan Hou, Jingyu Zhang, Tianxing He +7
Existing watermarking algorithms are vulnerable to paraphrase attacks because of their token-level design. To address this issue, we propose SemStamp, a robust sentence-level seman…