17 papers
FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents
Xianfu Cheng, Shiwei Zhang, Jiyu Zhao +10
Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. Howeve…
Persona Non Grata: LLM Persona-Driven Generations in MCQA are Unstable in Distinct Dimensions
César Guerra-Solano, Xiang Lorraine Li
Persona-driven generations (PDGs) have seen prolific use in research and industry applications, where a large language model (LLM) takes on a 'persona' while completing some task.…
SAGE: Scalable AI Governance & Evaluation
Benjamin Le, Xueying Lu, Nick Stern +17
Evaluating relevance in large-scale search systems is fundamentally constrained by the governance gap between nuanced, resource-constrained human oversight and the high-throughput…
Embodied Task Planning via Graph-Informed Action Generation with Large Language Models
Xiang Li, Ning Yan, Masood Mortazavi
While Large Language Models (LLMs) have demonstrated strong zero-shot reasoning capabilities, their deployment as embodied agents still faces fundamental challenges in long-horizon…
GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models
Zuyao Xu, Yuqi Qiu, Lu Sun +14
Citations provide the basis for trusting scientific claims; when they are invalid or fabricated, this trust collapses. With the advent of Large Language Models (LLMs), this risk ha…
Evaluating an evidence-guided reinforcement learning framework in aligning light-parameter large language models with decision-making cognition in psychiatric clinical reasoning
Xinxin Lin, Guangxin Dai, Yi Zhong +20
Large language models (LLMs) hold transformative potential for medical decision support yet their application in psychiatry remains constrained by hallucinations and superficial re…