212 citations · 323 across the 12 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2023
Proximal Policy Optimization Actual Combat: Manipulating Output Tokenizer Length
Miao Fan, Chen Hu, Shuchang Zhou
The Reinforcement Learning from Human Feedback (RLHF) plays a pivotal role in shaping the impact of large language models (LLMs), contributing significantly to controlling output t…
cs.AI2022★ 2 cited
ML4CO-KIDA: Knowledge Inheritance in Dataset Aggregation
Zixuan Cao, Yang Xu, Zhewei Huang +1
The Machine Learning for Combinatorial Optimization (ML4CO) NeurIPS 2021 competition aims to improve state-of-the-art combinatorial optimization solvers by replacing key heuristic…