From the 1 of 12 linked papers with an AI index.
12 papers
SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
Yang Zhou, Zixuan Huang, Sunzhu Li +10
The paper presents SpatialCLI, a framework that teaches vision-language models to use specialist visual tools for spatial reasoning and then internalize those capabilities, dramati…
EvA: An Evidence-First Audio Understanding Paradigm for LALMs
Xinyuan Xie, Shunian Chen, Zhiheng Liu +4
Large Audio Language Models (LALMs) still struggle in complex acoustic scenes because they often fail to preserve task-relevant acoustic evidence before reasoning begins. We identi…
Do Phone-Use Agents Respect Your Privacy?
Zhengyang Tang, Ke Ji, Xidong Wang +19
We study whether phone-use agents respect privacy while completing benign mobile tasks. This question has remained hard to answer because privacy-compliant behavior is not operatio…
From Lossy to Verified: A Provenance-Aware Tiered Memory for Agents
Qiming Zhu, Shunian Chen, Rui Yu +2
Long-horizon agents often compress interaction histories into write-time summaries. This creates a fundamental write-before-query barrier: compression decisions are made before the…
LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture
Xidong Wang, Dingjie Song, Shunian Chen +5
Expanding the long-context capabilities of Multi-modal Large Language Models~(MLLMs) is critical for advancing video understanding and high-resolution image analysis. Achieving thi…
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
Dingjie Song, Sicheng Lai, Mingxuan Wang +3
The rapid advancement of multimodal large language models (MLLMs) has significantly enhanced performance across benchmarks. However, data contamination-unintentional memorization o…