5 papers
SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model
Xiaoxin Lu, Ranran Haoran Zhang, Rui Zhang
Large language models (LLMs) are increasingly deployed as planners for autonomous agents in household environments. While existing benchmarks evaluate whether LLM-generated plans e…
Bridging the Know-Act Gap via Task-Level Autoregressive Reasoning
Jihyun Janice Ahn, Ryo Kamoi, Berk Atil +34
LLMs often generate seemingly valid answers to flawed or ill-posed inputs. This is not due to missing knowledge: under discriminative prompting, the same models can mostly identify…
Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation
Xiaoxin Lu, Ranran Haoran Zhang, Yusen Zhang +1
People get informed of a daily task plan through diverse media involving both texts and images. However, most prior research only focuses on LLM's capability of textual plan genera…
AAAR-1.0: Assessing AI's Potential to Assist Research
Renze Lou, Hanzi Xu, Sijia Wang +15
Numerous studies have assessed the proficiency of AI systems, particularly large language models (LLMs), in facilitating everyday tasks such as email writing, question answering, a…
HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?
Yusen Zhang, Wenliang Zheng, Aashrith Madasu +14
High-resolution image (HRI) understanding aims to process images with a large number of pixels, such as pathological images and agricultural aerial images, both of which can exceed…