From the 1 of 13 linked papers with an AI index.
13 papers
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
Shawn Li, Wei Yang, Jike Zhong +11
The paper introduces JigShape, a benchmark of interlocking jigsaw puzzles designed to test visual‑geometric reasoning in vision‑language models, and shows that current zero‑shot an…
Are Production Cloud Skills Adequately Tested? Measuring and Governing Skill Test Adequacy in Practice
Haotian Si, Junyi Chen, Shuyang Yu +5
Cloud platforms increasingly deliver reusable Cloud Skills that guide AI agents through multi-step resource operations, user choices, validation, and recovery. Existing Skill evalu…
RaMem: Contextual Reinstatement for Long-term Agentic Memory
Wei Yang, Bryce Kan, Shixuan Li +5
Long-term memory has become increasingly important for LLM agents that operate across extended interactions and evolving task contexts. Recent memory systems have made past experie…
Memory Retrieval for Changing Preferences
Yuehan Qin, Li Li, Linxin Song +4
Long-context dialogue systems must decide both when to access memory and which parts of the interaction history are relevant. Existing approaches typically rely on heuristic retrie…
Can Subgraph Explanations Be Weaponized to Steal Graph Neural Networks?
Ojas Nimase, Jiate Li, Yue Zhao +1
Graph Machine Learning as a Service (GMLaaS) platforms increasingly implement explainability interfaces to meet regulatory transparency requirements. However, this transparency cre…
"Someone Hid It": Query-Agnostic Black-Box Attacks on LLM-Based Retrieval
Jiate Li, Defu Cao, Li Li +8
Large language models (LLMs) have been serving as effective backbones for retrieval systems, including Retrieval-Augmentation-Generation (RAG), Dense Information Retriever (IR), an…