21 citations · 85 across the 29 of their papers we have counts for
29 papers
TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction
Jie Gong, Maowei Jiang, Zhiwei Liu +6
Large language models (LLMs) are increasingly used to make predictions from numerical time-series histories and textual events. Yet accuracy alone cannot reveal whether correct ans…
ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors
Jie Gong, Maowei Jiang, Zhiwei Liu +14
Conversational investment advisors influence not only what users know, but also how they make subsequent decisions as market conditions evolve. Existing evaluations primarily asses…
Not All Tokens Are Equal: Inflation-Aware Routing for Agentic LLM Systems
Heming Fu, Shan Lin, Qianqian Xie +1
When a language model fails to answer a query on the first attempt, an agentic system retries, consuming additional tokens each time. This retry overhead creates a gap between what…
Human or LLM as Standardized Patients? A Comparative Study for Medical Education
Bingquan Zhang, Xiaoxiao Liu, Yuchi Wang +3
Standardized patients (SPs) are indispensable for clinical skills training but remain expensive and difficult to scale. Although large language model (LLM)-based virtual standardiz…
DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation
Enze Zhang, Jiaying Wang, Mengxi Xiao +7
Large language models (LLMs) have substantially advanced machine translation (MT), yet their effectiveness in translating web novels remains unclear. Existing benchmarks rely on su…
MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application
Xueqing Peng, Lingfei Qian, Yan Wang +44
Real-world financial analysis involves information across multiple languages and modalities, from reports and news to scanned filings and meeting recordings. Yet most existing eval…