3 citations · 6 across the 13 of their papers we have counts for
Showing 2026 · cs.AIShow all
2 papers · 2 filters
cs.AI2026
TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction
Jie Gong, Maowei Jiang, Zhiwei Liu +6
Large language models (LLMs) are increasingly used to make predictions from numerical time-series histories and textual events. Yet accuracy alone cannot reveal whether correct ans…
cs.AI2026
Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation
Yan Wang, Yi Han, Lingfei Qian +11
Most recommendation benchmarks evaluate how well a model imitates user behavior. In financial advisory, however, observed actions can be noisy or short-sighted under market volatil…