8 papers
ToolVision: Learning When and How to Use Visual Tools with Capability-Aligned Supervision
Delin Mao, Chenghao Sun, Jingwei Song +2
Thinking with images allows a multimodal model to compensate for limited perception by invoking visual tools through code. Yet the prevailing SFT-then-RL recipe creates a different…
MemPrism: Task-Conditioned Relational Memory Views for Long-Horizon Agents
Zhisheng Chen, Bingfan Zeng, Bangde Cao +8
Long-horizon agents rely on memory to reuse experiences, yet existing memory systems often assume that evidence can be directly consumed through a fixed representation. This leads…
PersonaTree: Structured Lifecycle Memory for Person Understanding in LLM Agents
Yubo Hou, Jingwei Song, Hongbo Zhang +4
Persistent LLM agents require memory representations that make the formation of person understanding explicit across long term interaction. Existing agent memory methods emphasize…
PolarMem: A Training-Free Polarized Latent Graph Memory for Verifiable Vision-Language Models
Zhisheng Chen, Tingyu Wu, Zijie Zhou +7
Memory is not merely a storage mechanism for intelligent systems, but a structure for organizing evidence and constraining belief. This is especially important for multimodal reaso…
MineEvolve: Self-Evolution with Accumulated Knowledge for Long-Horizon Embodied Minecraft Agents
Zhengwei Xie, Zhisheng Chen, Ziyan Weng +7
Long-horizon embodied intelligence requires agents to improve through interaction, not merely to execute plans generated from static goals. A central challenge is therefore to tran…
Teaching Large Reasoning Models Effective Reflection
Hanbin Wang, Jingwei Song, Jinpeng Li +5
Large Reasoning Models (LRMs) have recently shown impressive performance on complex reasoning tasks, often by engaging in self-reflective behaviors such as self-critique and backtr…