15 papers
MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft
Tianjie Ju, Yueqing Sun, Zheng Wu +7
Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to sustain exploration in dynamic op…
HTMLCure: Turning Browser Experience into State Guided Repair for Interactive HTML
Jiajun Wu, Jian Yang, Tuney Zheng +4
LLMs can now produce full HTML pages, but many of those pages are only superficially correct: they render once, then fail under scroll, hover, click, resize, or gameplay. Evaluatio…
Beyond Fine-Tuning: In-Context Learning and Chain-of-Thought for Reasoned Distractor Generation
Elaf Alhazmi, Quan Z. Sheng, Wei Emma Zhang
Distractor generation (DG) remains a labor-intensive task that still significantly depends on domain experts. The task focuses on generating plausible yet incorrect options, known…
From Imitation to Discrimination: Progressive Curriculum Learning for Robust Web Navigation
Chuang Peng, Wei Zhang, Renshuai Tao +2
Text-based web agents offer computational efficiency for autonomous web navigation, yet developing robust agents remains challenging due to the noisy and heterogeneous nature of re…
SAGE: A Service Agent Graph-guided Evaluation Benchmark
Ling Shi, Yuqin Dai, Ziyin Wang +7
The development of Large Language Models (LLMs) has catalyzed automation in customer service, yet benchmarking their performance remains challenging. Existing benchmarks predominan…
IW-Bench: Evaluating Large Multimodal Models for Converting Image-to-Web
Hongcheng Guo, Wei Zhang, Junhao Chen +9
Recently advancements in large multimodal models have led to significant strides in image comprehension capabilities. Despite these advancements, there is a lack of the robust benc…