2 citations · 6 across the 30 of their papers we have counts for
9 papers · 1 filter
A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis
Fan Ma, Mauro Giuffrè, Donald Wright +12
Diagnostic error is a major threat to patient safety, yet current large language model (LLM) systems often treat diagnosis as a one-shot prediction task, lacking safeguards against…
Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation
Yuyang Dai, Xueqing Peng, Lingfei Qian +1
Evaluating the decision-making capabilities of large language models (LLMs) is a growing research priority, yet existing benchmarks focus on isolated cognitive tasks such as reason…
AUDITFLOW: Executable Symbolic Environments for Structured Financial Reporting Verification
Yan Wang, Xuguang Ai, Jaisal Patel +7
Structured financial audit verification is difficult for language-model agents because correctness depends on structured evidence rather than text alone. A model must link reported…
Herculean: An Agentic Benchmark for Financial Intelligence
Xueqing Peng, Zhuohan Xie, Yupeng Cao +60
As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…
Foundation Models to Unlock Real-World Evidence from Nationwide Medical Claims
Fan Ma, Yuntian Liu, Xiang Lan +22
Evidence derived from large-scale real-world data (RWD) is increasingly informing regulatory evaluation and healthcare decision-making. Administrative claims provide population-sca…
Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading
Polydoros Giannouris, Yuechen Jiang, Lingfei Qian +5
Many sequential decision-making problems exhibit hierarchical structure, where high-level semantic choices constrain downstream actions and feedback is delayed and ambiguous. Learn…