5 papers · 1 filter
LiAuto-GeoX: Efficient Grounded Driving Transformer
Jiawei Lian, Haoyi Sun, Yang Wu +8
Dense 3D reconstruction has demonstrated immense potential for spatial understanding, yet its viability as a real-time, onboard representation for autonomous driving remains an ope…
ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models
Kaiwen Xue, Tao Wei, Guoxin Zhang +5
Multimodal large language models (MLLMs) have shown strong potential as embodied agents, yet embodied geo-localization remains underexplored due to the lack of fine-grained evaluat…
CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding
Lihao Zheng, Zhenwei Shao, Yu Zhou +5
Although Multimodal Large Language Models (MLLMs) have advanced rapidly, they still face notable challenges in fine-grained multi-image understanding, often exhibiting spatial hall…
StreamingClaw Technical Report
Jiawei Chen, Zhe Chen, Chaoqun Du +21
Emerging applications such as embodied intelligence, AI hardware, autonomous driving, and intelligent cockpits rely on a real-time perception-decision-action closed loop, posing st…
MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning
Lihao Zheng, Jiawei Chen, Xintian Shen +2
Multi-image reasoning and grounding require understanding complex cross-image relationships at both object levels and image levels. Current Large Visual Language Models (LVLMs) fac…