2 papers
cs.CL2025
Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples
Chengqian Gao, Haonan Li, Liu Liu +3
The alignment of large language models (LLMs) often assumes that using more clean data yields better outcomes, overlooking the match between model capacity and example difficulty.…
cs.NE2024
Hard-Thresholding Meets Evolution Strategies in Reinforcement Learning
Chengqian Gao, William de Vazelhes, Hualin Zhang +2
Evolution Strategies (ES) have emerged as a competitive alternative for model-free reinforcement learning, showcasing exemplary performance in tasks like Mujoco and Atari. Notably,…