6 papers
Mint-Agent: Introducing Finance-Native Agentic Foundation Models
Agent Team, B. Zhang, Gavin Zhang +8
Financial agents must do more than recall domain knowledge: they must be both reliable, executing precise operations over grounded evidence, and executive, sustaining long-horizon…
HalluTracer: Hallucination Detection via Depth-Averaging Truth Signals
Zhihao Guo, Zonghan Wu, Huan Huo +6
Even well-aligned large language models confidently generate factually incorrect text, making hallucination a persistent reliability risk in high-stakes deployments. These models n…
TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning
Yaxuan Kong, Qingren Yao, Yuqi Nie +7
Time series data inform critical decisions across many real-world domains. While large language model (LLM) agents can analyze data through natural language and tools, it remains u…
SkillBrew: Multi-Objective Curation of Skill Banks for LLM Agents
Wentao Hu, Zhendong Chu, Yiming Zhang +6
Retrieval-augmented LLM agents increasingly rely on curated skill banks: collections of reusable textual principles that guide decision making on complex tasks. Existing approaches…
MIRL: Mutual Information-Guided Reinforcement Learning for Vision-Language Models
Yin Zhang, Jiaxuan Zhao, Zonghan Wu +5
Vision-Language Models (VLMs) frequently suffer from visual perception errors and hallucinations that compromise answer accuracy in complex reasoning tasks. Reinforcement Learning…
Towards Competent AI for Fundamental Analysis in Finance: A Benchmark Dataset and Evaluation
Zonghan Wu, Congyuan Zou, Junlin Wang +3
Generative AI, particularly large language models (LLMs), is beginning to transform the financial industry by automating tasks and helping to make sense of complex financial inform…