3 citations · 9 across the 5 of their papers we have counts for
5 papers
Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards
Wei Shen, Xiaoying Zhang, Yuanshun Yao +3
Reinforcement learning from human feedback (RLHF) is the mainstream paradigm used to align large language models (LLMs) with human preferences. Yet existing RLHF heavily relies on…
Learning to Watermark LLM-generated Text via Reinforcement Learning
Xiaojun Xu, Yuanshun Yao, Yang Liu
We study how to watermark LLM outputs, i.e. embedding algorithmically detectable signals into LLM-generated text to track misuse. Unlike the current mainstream methods that work wi…
Multi/Single-stage structured zero-gradient-sum approach for prescribed-time optimization
Shuaiyu Zhou, Yiheng Wei, Jinde Cao +1
Prescribed-time convergence mechanism has become a prominent research focus in the current field of optimization and control due to its ability to precisely control the target comp…
An Empirical Study of Malicious Code In PyPI Ecosystem
Wenbo Guo, Zhengzi Xu, Chengwei Liu +3
PyPI provides a convenient and accessible package management platform to developers, enabling them to quickly implement specific functions and improve work efficiency. However, the…
SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-training
Hong Yan, Yang Liu, Yushen Wei +3
Skeleton sequence representation learning has shown great advantages for action recognition due to its promising ability to model human joints and topology. However, the current me…