6 papers
Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use
Song-Lin Lv, Weiming Wu, Rui Zhu +2
While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, to…
MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs
Ziqiao Shang, Lingyue Ge, Ling-Yue Ge +12
Systematically evaluating Multimodal Large Language Models (MLLMs) is essential for advancing Artificial General Intelligence (AGI). Yet existing benchmarks remain inadequate for r…
TabularMath: Understanding Math Reasoning over Tables with Large Language Models
Shi-Yu Tian, Zhi Zhou, Wei Dong +5
Mathematical reasoning has long been a key benchmark for evaluating large language models. Although substantial progress has been made on math word problems, the need for reasoning…
LifeBench: A Benchmark for Long-Horizon Multi-Source Memory
Zihao Cheng, Weixin Wang, Yu Zhao +15
Long-term memory is fundamental for personalized agents capable of accumulating knowledge, reasoning over user experiences, and adapting across time. However, existing memory bench…
TabFSBench: Tabular Benchmark for Feature Shifts in Open Environments
Zi-Jian Cheng, Zi-Yi Jia, Zhi Zhou +2
Tabular data is widely utilized in various machine learning tasks. Current tabular learning research predominantly focuses on closed environments, while in real-world applications,…
Realistic Evaluation of TabPFN v2 in Open Environments
Zi-Jian Cheng, Zi-Yi Jia, Zhi Zhou +2
Tabular data, owing to its ubiquitous presence in real-world domains, has garnered significant attention in machine learning research. While tree-based models have long dominated t…