From the 1 of 29 linked papers with an AI index.
29 papers
NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing
Ming Yang, Zhi Zhou, Shi-Yu Tian +3
NeSy-Route is a large-scale neuro‑symbolic benchmark that provides automatically generated, constrained route‑planning tasks for remote‑sensing images, together with optimal soluti…
Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use
Song-Lin Lv, Weiming Wu, Rui Zhu +2
While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, to…
TriViewBench: Controlled Complexity Scaling for Multi-View Structural Reasoning in MLLMs
Yu-Yang Chen, Lan-Zhe Guo
Multimodal Large Language Models (MLLMs) demonstrate strong performance on standard visual question answering benchmarks, yet their scalability under controlled structural complexi…
On the Learnability of Test-Time Adaptation: A Recovery Complexity Perspective
Zhi Zhou, Ming Yang, Shi-Yu Tian +3
Test-time adaptation (TTA) aims to adapt models to maintain reliable performance on non-stationary test streams without requiring labeled data. Despite its empirical success, the l…
Roles with Rails: Contract-Preserving Role Evolution in Multi-Agent Structured Reasoning
Ling-Yue Ge, Lan-Zhe Guo
Role-based LLM multi-agent systems need adaptive role pools, yet adapting such systems is not merely a matter of prompt optimization: roles often carry structural obligations, incl…
VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning
Zi-Yi Jia, Zi-Jian Cheng, Xin-Yue Zhang +4
Multi-model learning has attracted great attention in visual-text tasks. However, visual-tabular data, which plays a pivotal role in high-stakes domains like healthcare and industr…