1 citations · 1 across the 11 of their papers we have counts for
5 papers · 1 filter
Information-Time Proximal Policy Optimization
Yongcheng Zeng, Xinyu Cui, Yan Song +9
RLVR has substantially improved the reasoning capabilities of LLMs. However, existing methods typically parameterize temporal progression in the Markov Decision Process by token-by…
Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search
Zhongwei Yu, Yan Song, Xue Yan +9
Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended hypothesis spaces, such as molecules, protein sequences, and c…
Memory-Driven Self-Improvement for Decision Making with Large Language Models
Xue Yan, Zijing Ou, Mengyue Yang +4
Large language models (LLMs) have emerged as effective action policies for sequential decision-making (SDM) tasks due to their extensive prior knowledge. However, this broad yet ge…
Efficient Reinforcement Learning with Large Language Model Priors
Xue Yan, Yan Song, Xidong Feng +4
In sequential decision-making (SDM) tasks, methods like reinforcement learning (RL) and heuristic search have made notable advances in specific cases. However, they often require e…
Ask more, know better: Reinforce-Learned Prompt Questions for Decision Making with Large Language Models
Xue Yan, Yan Song, Xinyu Cui +4
Large language models (LLMs) demonstrate their promise in tackling complicated practical challenges by combining action-based policies with chain of thought (CoT) reasoning. Having…