7 papers
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
Tianyu Xu, Jiawei Chen, Jiazhao Zhang +5
Visual navigation policy is widely regarded as a promising direction, as it mimics humans by using egocentric visual observations for navigation. However, optical information of vi…
CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding
Lihao Zheng, Zhenwei Shao, Yu Zhou +5
Although Multimodal Large Language Models (MLLMs) have advanced rapidly, they still face notable challenges in fine-grained multi-image understanding, often exhibiting spatial hall…
StreamingClaw Technical Report
Jiawei Chen, Zhe Chen, Chaoqun Du +21
Emerging applications such as embodied intelligence, AI hardware, autonomous driving, and intelligent cockpits rely on a real-time perception-decision-action closed loop, posing st…
SPAN-Nav: Generalized Spatial Awareness for Versatile Vision-Language Navigation
Jiahang Liu, Tianyu Xu, Jiawei Chen +9
Recent embodied navigation approaches leveraging Vision-Language Models (VLMs) demonstrate strong generalization in versatile Vision-Language Navigation (VLN). However, reliable pa…
Evolving from Tool User to Creator via Training-Free Experience Reuse in Multimodal Reasoning
Xintian Shen, Jiawei Chen, Lihao Zheng +3
Existing Tool-Integrated Reasoning (TIR) models have effectively extended the question-answering capabilities of LLMs by incorporating external tools. However, real-world scenarios…
MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning
Jiawei Chen, Xintian Shen, Lihao Zheng +43
Traditional workflow-based agents exhibit limited intelligence when addressing real-world problems requiring tool invocation. Tool-integrated reasoning (TIR) agents capable of auto…