6 papers
SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning
Keyang Zhong, Kuo Wang, Peng Liu +5
Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging current multimodal models with limited co…
X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction
Xiaoming Ren, Ru Zhen, Chao Li +11
Inspired by the development of OpenClaw, there is a growing demand for mobile-based personal agents capable of handling complex and intuitive interactions. In this technical report…
AndesVL Technical Report: An Efficient Mobile-side Multimodal Large Language Model
Zhiwei Jin, Xiaohui Song, Nan Wang +36
In recent years, while cloud-based MLLMs such as QwenVL, InternVL, GPT-4o, Gemini, and Claude Sonnet have demonstrated outstanding performance with enormous model sizes reaching hu…
Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference
Kuo Wang, Quanlong Zheng, Junlin Xie +6
Video Multimodal Large Language Models~(Video-MLLM) have achieved remarkable advancements in video understanding tasks. However, constrained by the context length limitation in the…
Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM
Peng Liu, Xiaoming Ren, Fengkai Liu +5
Recent advancements in image-to-video (I2V) generation have shown promising performance in conventional scenarios. However, these methods still encounter significant challenges whe…
H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding
Qi Wu, Quanlong Zheng, Yanhao Zhang +8
With the rapid development of multimodal models, the demand for assessing video understanding capabilities has been steadily increasing. However, existing benchmarks for evaluating…