From the 1 of 4 linked papers with an AI index.
4 papers
Efficient Test-Time Optimization for Multi-Agent Proof Autoformalization
Tian-Shuo Liu, Shiyuan Zhang, Zijie Geng +5
The paper introduces ToMap, a multi‑agent system that treats proof autoformalization as a Decomposer‑Formalizer‑Prover pipeline and concentrates test‑time optimization on improving…
Off-Policy Value-Based Reinforcement Learning for Large Language Models
Peng-Yuan Wang, Ziniu Li, Tian Xu +8
Improving data utilization efficiency is critical for scaling reinforcement learning (RL) for long-horizon tasks where generating trajectories is expensive. However, the dominant R…
A Survey on Large Language Models for Mathematical Reasoning
Peng-Yuan Wang, Tian-Shuo Liu, Chenyang Wang +8
Mathematical reasoning has long represented one of the most fundamental and challenging frontiers in artificial intelligence research. In recent years, large language models (LLMs)…
Controlling Large Language Model with Latent Actions
Chengxing Jia, Ziniu Li, Pengyuan Wang +4
Adapting Large Language Models (LLMs) to downstream tasks using Reinforcement Learning (RL) has proven to be an effective approach. However, LLMs do not inherently define the struc…