4 papers · 1 filter
AUDITFLOW: Executable Symbolic Environments for Structured Financial Reporting Verification
Yan Wang, Xuguang Ai, Jaisal Patel +7
Structured financial audit verification is difficult for language-model agents because correctness depends on structured evidence rather than text alone. A model must link reported…
Herculean: An Agentic Benchmark for Financial Intelligence
Xueqing Peng, Zhuohan Xie, Yupeng Cao +60
As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…
StaRPO: Stability-Augmented Reinforcement Policy Optimization
Jinghan Zhang, Fengran Mo, Tharindu Cyril Weerasooriya +5
Reinforcement learning (RL) is effective in enhancing the accuracy of large language models in complex reasoning tasks. Existing RL policy optimization frameworks rely on final-ans…
Entropy-based Exploration Conduction for Multi-step Reasoning
Jinghan Zhang, Xiting Wang, Fengran Mo +3
Multi-step processes via large language models (LLMs) have proven effective for solving complex reasoning tasks. However, the depth of exploration of the reasoning procedure can si…