4 papers
AgentRewind: Recoverable Execution for Long-Horizon LLM Agents
Yu Zhuang, Kefei Chen, Yitong Duan +3
Many real-world tasks require LLM agents to interact with their environments over long execution horizons. Errors that occur early in execution may propagate through both the agent…
M3: A State-Event Generative Foundation Model for Market Microstructure Dynamics
Yanzhi Zhang, Yu Ma, Yilin Cheng +2
Market microstructure simulation aims to model how liquidity, prices, and order flow evolve in electronic financial markets. Since market data reveal only one realized trajectory,…
Population-Evolve: a Parallel Sampling and Evolutionary Method for LLM Math Reasoning
Yanzhi Zhang, Yitong Duan, Zhaoxi Zhang +2
Test-time scaling has emerged as a promising direction for enhancing the reasoning capabilities of Large Language Models in last few years. In this work, we propose Population-Evol…
No Free Lunch: Rethinking Internal Feedback for LLM Reasoning
Yanzhi Zhang, Zhaoxi Zhang, Haoxiang Guan +6
Reinforcement learning has emerged as a powerful paradigm for post-training large language models (LLMs) to improve reasoning. Approaches like Reinforcement Learning from Human Fee…