4 papers
TemporalBench: A Benchmark for Evaluating LLM-Based Agents on Contextual and Event-Informed Time Series Tasks
Muyan Weng, Defu Cao, Wei Yang +2
It is unclear whether strong forecasting performance reflects genuine temporal understanding or the ability to reason under contextual and event-driven conditions. We introduce Tem…
MOSAIC: Modular Foundation Models for Assistive and Interactive Cooking
Huaxiaoyue Wang, Kushal Kedia, Juntao Ren +14
We present MOSAIC, a modular architecture for coordinating multiple robots to (a) interact with users using natural language and (b) manipulate an open vocabulary of everyday objec…
Advancing Conversational Diagnostic AI with Multimodal Reasoning
Khaled Saab, Jan Freyberg, Chunjong Park +33
Large Language Models (LLMs) have demonstrated great potential for conducting diagnostic conversations but evaluation has been largely limited to language-only interactions, deviat…
Attribute Diversity Determines the Systematicity Gap in VQA
Ian Berlot-Attwell, Kumar Krishna Agrawal, A. Michael Carrell +2
Although modern neural networks often generalize to new combinations of familiar concepts, the conditions that enable such compositionality have long been an open question. In this…