collaborators

6 papers

cs.AI2026

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use

Song-Lin Lv, Weiming Wu, Rui Zhu +2

While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, to…

cs.LG2026

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs

Ziqiao Shang, Lingyue Ge, Ling-Yue Ge +12

Systematically evaluating Multimodal Large Language Models (MLLMs) is essential for advancing Artificial General Intelligence (AGI). Yet existing benchmarks remain inadequate for r…

cs.AI2026

TabularMath: Understanding Math Reasoning over Tables with Large Language Models

Shi-Yu Tian, Zhi Zhou, Wei Dong +5

Mathematical reasoning has long been a key benchmark for evaluating large language models. Although substantial progress has been made on math word problems, the need for reasoning…

cs.AI2026

LifeBench: A Benchmark for Long-Horizon Multi-Source Memory

Zihao Cheng, Weixin Wang, Yu Zhao +15

Long-term memory is fundamental for personalized agents capable of accumulating knowledge, reasoning over user experiences, and adapting across time. However, existing memory bench…

cs.LG2025

TabFSBench: Tabular Benchmark for Feature Shifts in Open Environments

Zi-Jian Cheng, Zi-Yi Jia, Zhi Zhou +2

Tabular data is widely utilized in various machine learning tasks. Current tabular learning research predominantly focuses on closed environments, while in real-world applications,…

cs.LG2025

Realistic Evaluation of TabPFN v2 in Open Environments

Zi-Jian Cheng, Zi-Yi Jia, Zhi Zhou +2

Tabular data, owing to its ubiquitous presence in real-world domains, has garnered significant attention in machine learning research. While tree-based models have long dominated t…