3 papers
cs.AI2026
Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents
Yuanhong Jiang, Jingjie Zou, Rui Jiang +5
Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundari…
cs.AI2026
BizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation
Xin Guo, Rongjunchen Zhang, Guilong Lu +4
Large language models are becoming increasingly significant in financial applications. Nevertheless, prevailing benchmarks are largely dependent on simulated or generic data, which…
cs.SE2026
SGCR: A Specification-Grounded Framework for Trustworthy LLM Code Review
Kai Wang, Bingcheng Mao, Shuai Jia +4
Automating code review with Large Language Models (LLMs) shows immense promise, yet practical adoption is hampered by their lack of reliability, context-awareness, and control. To…