20 papers
Herculean: An Agentic Benchmark for Financial Intelligence
Xueqing Peng, Zhuohan Xie, Yupeng Cao +60
As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…
LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment
Lingyao Li, Deyi Li, Chen Chen +9
Large language models (LLMs) are increasingly deployed across healthcare applications, including clinical documentation, diagnostic reasoning, medicine recommendation, and medical…
Concordia: Self-Improving Synthetic Tables for Federated LLMs
Jimin Huang, Duanyu Feng, Nuo Chen +8
Federated learning (FL) enables training large language models (LLMs) without sharing raw data, but adapting LLMs under strict data isolation and non-IID client distributions remai…
CXR-LT 2026 Challenge: Multi-Center Long-Tailed and Zero Shot Chest X-ray Classification
Hexin Dong, Yi Lin, Pengyu Zhou +25
Chest X-ray (CXR) interpretation is hindered by the long-tailed distribution of pathologies and the open-world nature of clinical environments. Existing benchmarks often rely on cl…
FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR
Yueru He, Xueqing Peng, Yupeng Cao +13
Recent progress in multimodal large language models (MLLMs) has substantially improved document understanding, yet strong optical character recognition (OCR) performance on surface…
PRIME: Prototype-Driven Multimodal Pretraining for Cancer Prognosis with Missing Modalities
Kai Yu, Shuang Zhou, Yiran Song +9
Multimodal self-supervised pretraining offers a promising route to cancer prognosis by integrating histopathology whole-slide images, gene expression, and pathology reports, yet mo…