5 papers
FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks
Yupeng Cao, Haohang Li, Weijin Liu +11
Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for long-horizon financial tasks. While existing…
Truth Neurons
Haohang Li, Yupeng Cao, Yangyang Yu +2
Despite their remarkable success and deployment across diverse workflows, language models sometimes produce untruthful responses. Our limited understanding of how truthfulness is m…
Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications
Jimin Huang, Mengxi Xiao, Dong Li +41
Financial LLMs hold promise for advancing financial tasks and domain-specific applications. However, they are limited by scarce corpora, weak multimodal capabilities, and narrow ev…
Replicating Human Social Perception in Generative AI: Evaluating the Valence-Dominance Model
Necdet Gurkan, Kimathi Njoki, Jordan W. Suchow
As artificial intelligence (AI) continues to advance--particularly in generative models--an open question is whether these systems can replicate foundational models of human social…
INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent
Haohang Li, Yupeng Cao, Yangyang Yu +12
Recent advancements have underscored the potential of large language model (LLM)-based agents in financial decision-making. Despite this progress, the field currently encounters tw…