17 papers
SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning
Keyang Zhong, Kuo Wang, Peng Liu +5
Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multimodal models with limited co…
PIGEON: VLM-Driven Object Navigation via Points of Interest Selection
Cheng Peng, Zhenzhe Zhang, Xiaobao Wei +7
Object navigation in unseen indoor environments requires agents to perform semantic search under partial observability. Vision-language models (VLMs) provide strong semantic-spatia…
MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging
Luyuan Zhang, Siyuan Li, Zedong Wang +7
Most visual tokenizers for image generation are bifurcated into two families with complementary limitations: continuous VAEs offer high-fidelity reconstruction but suffer from dens…
X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction
Xiaoming Ren, Ru Zhen, Chao Li +11
Inspired by the development of OpenClaw, there is a growing demand for mobile-based personal agents capable of handling complex and intuitive interactions. In this technical report…
OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning
Zhijia Liang, Jiaming Li, Weikai Chen +3
Streaming video reasoning requires models to operate in a setting where history grows without bound while meaningful evidence remains scarce. In such a landscape, relevant signal i…
Thinking in Streaming Video
Zikang Liu, Longteng Guo, Handong Li +7
Real-time understanding of continuous video streams is essential for interactive assistants and multimodal agents operating in dynamic environments. However, most existing video re…