3 papers
cs.AI2026
BizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation
Xin Guo, Rongjunchen Zhang, Guilong Lu +4
Large language models are becoming increasingly significant in financial applications. Nevertheless, prevailing benchmarks are largely dependent on simulated or generic data, which…
cs.CL2026
Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
Yuhan Guo, Cong Guo, Aiwen Sun +12
Multimodal large-scale models have significantly advanced the development of web agents, enabling perception and interaction with digital environments akin to human cognition. In t…
cs.AI2025
BizFinBench: A Business-Driven Real-World Financial Benchmark for Evaluating LLMs
Guilong Lu, Xuntao Guo, Rongjunchen Zhang +2
Large language models excel in general tasks, yet assessing their reliability in logic-heavy, precision-critical domains like finance, law, and healthcare remains challenging. To a…