2 papers
cs.AI2026
TemporalBench: A Benchmark for Evaluating LLM-Based Agents on Contextual and Event-Informed Time Series Tasks
Muyan Weng, Defu Cao, Wei Yang +2
It is unclear whether strong forecasting performance reflects genuine temporal understanding or the ability to reason under contextual and event-driven conditions. We introduce Tem…
cs.CL2025
Advancing Conversational Diagnostic AI with Multimodal Reasoning
Khaled Saab, Jan Freyberg, Chunjong Park +33
Large Language Models (LLMs) have demonstrated great potential for conducting diagnostic conversations but evaluation has been largely limited to language-only interactions, deviat…