works on

From the 1 of 29 linked papers with an AI index.

collaborators

29 papers

cs.AI2026

NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing

Ming Yang, Zhi Zhou, Shi-Yu Tian +3

NeSy-Route is a large-scale neuro‑symbolic benchmark that provides automatically generated, constrained route‑planning tasks for remote‑sensing images, together with optimal soluti…

cs.AI2026

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use

Song-Lin Lv, Weiming Wu, Rui Zhu +2

While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, to…

cs.CV2026

TriViewBench: Controlled Complexity Scaling for Multi-View Structural Reasoning in MLLMs

Yu-Yang Chen, Lan-Zhe Guo

Multimodal Large Language Models (MLLMs) demonstrate strong performance on standard visual question answering benchmarks, yet their scalability under controlled structural complexi…

cs.LG2026

On the Learnability of Test-Time Adaptation: A Recovery Complexity Perspective

Zhi Zhou, Ming Yang, Shi-Yu Tian +3

Test-time adaptation (TTA) aims to adapt models to maintain reliable performance on non-stationary test streams without requiring labeled data. Despite its empirical success, the l…

cs.CL2026

Roles with Rails: Contract-Preserving Role Evolution in Multi-Agent Structured Reasoning

Ling-Yue Ge, Lan-Zhe Guo

Role-based LLM multi-agent systems need adaptive role pools, yet adapting such systems is not merely a matter of prompt optimization: roles often carry structural obligations, incl…

cs.CV2026

VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning

Zi-Yi Jia, Zi-Jian Cheng, Xin-Yue Zhang +4

Multi-model learning has attracted great attention in visual-text tasks. However, visual-tabular data, which plays a pivotal role in high-stakes domains like healthcare and industr…