activity
20242026
collaborators

32 papers

cs.CL2026

Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval

Linhai Ma, Ethan F. Wei, Xueqing Peng +3

Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however, the input is indirect evidence, such as a table cell whose me…

cs.AI2026

Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?

Yuyang Dai, Xueqing Peng, Yuxia Wang +2

Large language models are increasingly applied as autonomous decision-making agents. However, in executive business decisions, existing benchmarks are limited to textonly settings.…

cs.CL2026

Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering

Zhuohan Xie, Xueqing Peng, Georgi Georgiev +18

FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and n…

cs.CL2026

Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering

Zhuohan Xie, Yuyang Dai, Rania Elbadry +18

FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the corr…

cs.CL2026

FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents

Muhammad Usman Safder, Ayesha Gull, Rania Elbadry +7

Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as "preserve capital" or "avoid speculative…

cs.AI2026

Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation

Yuyang Dai, Xueqing Peng, Lingfei Qian +1

Evaluating the decision-making capabilities of large language models (LLMs) is a growing research priority, yet existing benchmarks focus on isolated cognitive tasks such as reason…