collaborators

5 papers

cs.AI2026

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales

Tianyu Liu, Allen Xin Wang, Antonia Panescu +30

AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood. Existing benchma…

cs.GT2026

Training Language Models for Bilateral Trade with Private Information

Dirk Bergemann, Soheil Ghili, Xinyang Hu +2

Bilateral bargaining under incomplete information provides a controlled testbed for evaluating large language model (LLM) agent capabilities. Bilateral trade demands individual rat…

cs.LG2025

Understanding In-context Learning of Addition via Activation Subspaces

Xinyan Hu, Kayo Yin, Michael I. Jordan +2

To perform few-shot learning, language models extract signals from a few input-label pairs, aggregate these into a learned prediction rule, and apply this rule to new inputs. How i…

cs.CV2025

FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging

Zichen Tang, Haihong E, Jiacheng Liu +18

We present FinMMR, a novel bilingual multimodal benchmark tailored to evaluate the reasoning capabilities of multimodal large language models (MLLMs) in financial numerical reasoni…

cs.CL2025

FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging

Zichen Tang, Haihong E, Ziyan Ma +10

We introduce FinanceReasoning, a novel benchmark designed to evaluate the reasoning capabilities of large reasoning models (LRMs) in financial numerical reasoning problems. Compare…