collaborators

14 papers

cs.CL2026

Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering

Zhuohan Xie, Xueqing Peng, Georgi Georgiev +18

FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and n…

cs.CL2026

Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering

Zhuohan Xie, Yuyang Dai, Rania Elbadry +18

FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the corr…

cs.HC2026

From Prediction to Intervention: Personalized Meal-Level Glucose Regulation via an LLM Agent

Mingyu Huang, Weiqing Min, Ying Jin +2

Personalized glucose regulation remains a central yet unresolved challenge in precision nutrition, as postprandial glucose response varies substantially across individuals. Existin…

cs.CL2026

FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents

Muhammad Usman Safder, Ayesha Gull, Rania Elbadry +7

Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as "preserve capital" or "avoid speculative…

cs.CL2026

EngTrace: A Symbolic Benchmark for Verifiable Process Supervision of Engineering Reasoning

Ayesha Gull, Muhammad Usman Safder, Rania Elbadry +4

Large Language Models (LLMs) are increasingly entering specialized, safety-critical engineering workflows governed by strict quantitative standards and immutable physical laws, mak…

cs.AI2026

Herculean: An Agentic Benchmark for Financial Intelligence

Xueqing Peng, Zhuohan Xie, Yupeng Cao +60

As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…