8 papers
LLM-Based World Models Can Make Decisions Solely, But Rigorous Evaluations are Needed
Chang Yang, Xinrun Wang, Junzhe Jiang +2
World model emerges as a key module in decision making, where MuZero and Dreamer achieve remarkable successes in complex tasks. Recent work leverages Large Language Models (LLMs) a…
Nondeterministic Polynomial-time Problem Challenge: An Ever-Scaling Reasoning Benchmark for LLMs
Chang Yang, Ruiyu Wang, Junzhe Jiang +9
Reasoning is the fundamental capability of large language models (LLMs). Due to the rapid progress of LLMs, there are two main issues of current benchmarks: i) these benchmarks can…
Offline-to-Online Reinforcement Learning with Classifier-Free Diffusion Generation
Xiao Huang, Xu Liu, Enze Zhang +2
Offline-to-online Reinforcement Learning (O2O RL) aims to perform online fine-tuning on an offline pre-trained policy to minimize costly online interactions. Existing work used off…
FinMaster: A Holistic Benchmark for Mastering Full-Pipeline Financial Workflows with LLMs
Junzhe Jiang, Chang Yang, Aixin Cui +6
Financial tasks are pivotal to global economic stability; however, their execution faces challenges including labor intensive processes, low error tolerance, data fragmentation, an…
Resolving Latency and Inventory Risk in Market Making with Reinforcement Learning
Junzhe Jiang, Chang Yang, Xinrun Wang +3
The latency of the exchanges in Market Making (MM) is inevitable due to hardware limitations, system processing times, delays in receiving data from exchanges, the time required fo…
In-Context Exploiter for Extensive-Form Games
Shuxin Li, Chang Yang, Youzhi Zhang +5
Nash equilibrium (NE) is a widely adopted solution concept in game theory due to its stability property. However, we observe that the NE strategy might not always yield the best re…