Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management
Yuxuan Zhao, Sijia Chen, Ningxin Su
Large language models (LLMs) have shown strong performance across diverse financial tasks, yet portfolio management (PM) remains poorly benchmarked. Existing benchmarks exhibit two…
cs.AI2026
EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation
Yi Liu, TingFeng Hui, Wei Zhang +4
Scalable AI agents training relies on interactive environments that faithfully simulate the consequences of agent actions. Manually crafted environments are expensive to build, bri…