32 papers
Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval
Linhai Ma, Ethan F. Wei, Xueqing Peng +3
Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however, the input is indirect evidence, such as a table cell whose me…
Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?
Yuyang Dai, Xueqing Peng, Yuxia Wang +2
Large language models are increasingly applied as autonomous decision-making agents. However, in executive business decisions, existing benchmarks are limited to textonly settings.…
Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering
Zhuohan Xie, Xueqing Peng, Georgi Georgiev +18
FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and n…
Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering
Zhuohan Xie, Yuyang Dai, Rania Elbadry +18
FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the corr…
FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents
Muhammad Usman Safder, Ayesha Gull, Rania Elbadry +7
Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as "preserve capital" or "avoid speculative…
Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation
Yuyang Dai, Xueqing Peng, Lingfei Qian +1
Evaluating the decision-making capabilities of large language models (LLMs) is a growing research priority, yet existing benchmarks focus on isolated cognitive tasks such as reason…