collaborators

6 papers

cs.AI2026

Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation

Yan Wang, Yi Han, Lingfei Qian +11

Most recommendation benchmarks evaluate how well a model imitates user behavior. In financial advisory, however, observed actions can be noisy or short-sighted under market volatil…

cs.CL2026

FinTagging: Benchmarking LLMs for Extracting and Structuring Financial Information

Yan Wang, Lingfei Qian, Xueqing Peng +18

Accurate interpretation of numerical data in financial reports is critical for markets and regulators. Although XBRL (eXtensible Business Reporting Language) provides a standard fo…

cs.AI2026

Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment

Yi Han, Yan Wang, Lingfei Qian +12

Large language model (LLM) agents are increasingly tested on complex tasks, but their ability to allocate scarce resources over long horizons remains unclear. Unlike reactive tasks…

cs.AI2025

Knowledge-Guided Adaptive Mixture of Experts for Precipitation Prediction

Chen Jiang, Kofi Osei, Sai Deepthi Yeddula +2

Accurate precipitation forecasting is indispensable in agriculture, disaster management, and sustainable strategies. However, predicting rainfall has been challenging due to the co…

cs.IR2025

OrdRankBen: A Novel Ranking Benchmark for Ordinal Relevance in NLP

Yan Wang, Lingfei Qian, Xueqing Peng +2

The evaluation of ranking tasks remains a significant challenge in natural language processing (NLP), particularly due to the lack of direct labels for results in real-world scenar…

cs.CL2025

Prompting a Weighting Mechanism into LLM-as-a-Judge in Two-Step: A Case Study

Wenwen Xie, Gray Gwizdz, Dongji Feng

While Large Language Models (LLMs) have emerged as promising tools for evaluating Natural Language Generation (NLG) tasks, their effectiveness is limited by their inability to appr…