collaborators

17 papers

cs.AI2026

SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning

Keyang Zhong, Kuo Wang, Peng Liu +5

Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multimodal models with limited co…

cs.RO2026

PIGEON: VLM-Driven Object Navigation via Points of Interest Selection

Cheng Peng, Zhenzhe Zhang, Xiaobao Wei +7

Object navigation in unseen indoor environments requires agents to perform semantic search under partial observability. Vision-language models (VLMs) provide strong semantic-spatia…

cs.CV2026

MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging

Luyuan Zhang, Siyuan Li, Zedong Wang +7

Most visual tokenizers for image generation are bifurcated into two families with complementary limitations: continuous VAEs offer high-fidelity reconstruction but suffer from dens…

cs.CV2026

X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction

Xiaoming Ren, Ru Zhen, Chao Li +11

Inspired by the development of OpenClaw, there is a growing demand for mobile-based personal agents capable of handling complex and intuitive interactions. In this technical report…

cs.CV2026

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning

Zhijia Liang, Jiaming Li, Weikai Chen +3

Streaming video reasoning requires models to operate in a setting where history grows without bound while meaningful evidence remains scarce. In such a landscape, relevant signal i…

cs.CV2026

Thinking in Streaming Video

Zikang Liu, Longteng Guo, Handong Li +7

Real-time understanding of continuous video streams is essential for interactive assistants and multimodal agents operating in dynamic environments. However, most existing video re…