works on

From the 1 of 6 linked papers with an AI index.

activity
20242026
collaborators

6 papers

cs.AI2026

Efficient Test-Time Optimization for Multi-Agent Proof Autoformalization

Tian-Shuo Liu, Shiyuan Zhang, Zijie Geng +5

The paper introduces ToMap, a multi‑agent system that treats proof autoformalization as a Decomposer‑Formalizer‑Prover pipeline and concentrates test‑time optimization on improving…

cs.LG2026

Off-Policy Value-Based Reinforcement Learning for Large Language Models

Peng-Yuan Wang, Ziniu Li, Tian Xu +8

Improving data utilization efficiency is critical for scaling reinforcement learning (RL) for long-horizon tasks where generating trajectories is expensive. However, the dominant R…

cs.AI2025

A Survey on Large Language Models for Mathematical Reasoning

Peng-Yuan Wang, Tian-Shuo Liu, Chenyang Wang +8

Mathematical reasoning has long represented one of the most fundamental and challenging frontiers in artificial intelligence research. In recent years, large language models (LLMs)…

cs.CL2025

Controlling Large Language Model with Latent Actions

Chengxing Jia, Ziniu Li, Pengyuan Wang +4

Adapting Large Language Models (LLMs) to downstream tasks using Reinforcement Learning (RL) has proven to be an effective approach. However, LLMs do not inherently define the struc…

cs.LG2024

WHALE: Towards Generalizable and Scalable World Models for Embodied Decision-making

Zhilong Zhang, Ruifeng Chen, Junyin Ye +8

World models play a crucial role in decision-making within embodied environments, enabling cost-free explorations that would otherwise be expensive in the real world. To facilitate…

cs.CL2024

BWArea Model: Learning World Model, Inverse Dynamics, and Policy for Controllable Language Generation

Chengxing Jia, Pengyuan Wang, Ziniu Li +4

Large language models (LLMs) have catalyzed a paradigm shift in natural language processing, yet their limited controllability poses a significant challenge for downstream applicat…