From the 1 of 6 linked papers with an AI index.
6 papers
Efficient Test-Time Optimization for Multi-Agent Proof Autoformalization
Tian-Shuo Liu, Shiyuan Zhang, Zijie Geng +5
The paper introduces ToMap, a multi‑agent system that treats proof autoformalization as a Decomposer‑Formalizer‑Prover pipeline and concentrates test‑time optimization on improving…
Off-Policy Value-Based Reinforcement Learning for Large Language Models
Peng-Yuan Wang, Ziniu Li, Tian Xu +8
Improving data utilization efficiency is critical for scaling reinforcement learning (RL) for long-horizon tasks where generating trajectories is expensive. However, the dominant R…
A Survey on Large Language Models for Mathematical Reasoning
Peng-Yuan Wang, Tian-Shuo Liu, Chenyang Wang +8
Mathematical reasoning has long represented one of the most fundamental and challenging frontiers in artificial intelligence research. In recent years, large language models (LLMs)…
Controlling Large Language Model with Latent Actions
Chengxing Jia, Ziniu Li, Pengyuan Wang +4
Adapting Large Language Models (LLMs) to downstream tasks using Reinforcement Learning (RL) has proven to be an effective approach. However, LLMs do not inherently define the struc…
WHALE: Towards Generalizable and Scalable World Models for Embodied Decision-making
Zhilong Zhang, Ruifeng Chen, Junyin Ye +8
World models play a crucial role in decision-making within embodied environments, enabling cost-free explorations that would otherwise be expensive in the real world. To facilitate…
BWArea Model: Learning World Model, Inverse Dynamics, and Policy for Controllable Language Generation
Chengxing Jia, Pengyuan Wang, Ziniu Li +4
Large language models (LLMs) have catalyzed a paradigm shift in natural language processing, yet their limited controllability poses a significant challenge for downstream applicat…