collaborators

6 papers

cs.AI2026

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks

Yupeng Cao, Haohang Li, Weijin Liu +11

Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for long-horizon financial tasks. While existing…

cs.CE2025

FinAudio: A Benchmark for Audio Large Language Models in Financial Applications

Yupeng Cao, Haohang Li, Yangyang Yu +10

Audio Large Language Models (AudioLLMs) have received widespread attention and have significantly improved performance on audio tasks such as conversation, audio understanding, and…

cs.AI2025

Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting

Shashidhar Reddy Javaji, Bhavul Gauri, Zining Zhu

Large language models (LLMs) are now used in multi-turn workflows, but we still lack a clear way to measure when iteration helps and when it hurts. We present an evaluation framewo…

cs.CL2025

Truth Neurons

Haohang Li, Yupeng Cao, Yangyang Yu +2

Despite their remarkable success and deployment across diverse workflows, language models sometimes produce untruthful responses. Our limited understanding of how truthfulness is m…

cs.LG2025

VERBA: Verbalizing Model Differences Using Large Language Models

Shravan Doda, Shashidhar Reddy Javaji, Zining Zhu

In the current machine learning landscape, we face a "model lake" phenomenon: Given a task, there is a proliferation of trained models with similar performances despite different b…

cs.CL2025

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim Evidence Reasoning

Shashidhar Reddy Javaji, Yupeng Cao, Haohang Li +3

Large language models (LLMs) are increasingly being used for complex research tasks such as literature review, idea generation, and scientific paper analysis, yet their ability to…