7 citations · 13 across the 14 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Is Crowdsourcing Breaking Your Bank? Cost-Effective Fine-Tuning of Pre-trained Language Models with Proximal Policy Optimization
Shuo Yang, Gjergji Kasneci
Wide usage of ChatGPT has highlighted the potential of reinforcement learning from human feedback. However, its training pipeline relies on manual ranking, a resource-intensive pro…
cs.CL2023★ 7 cited
Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Shuo Yang, Wei-Lin Chiang, Lianmin Zheng +2
Large language models are increasingly trained on all the data ever produced by humans. Many have raised concerns about the trustworthiness of public benchmarks due to potential co…