1 citations · 1 across the 4 of their papers we have counts for
4 papers
Geometry-Guided 3D Visual Token Pruning for Video-Language Models
Han Li, Zehao Huang, Jiahui Fu +2
Multimodal large language models have demonstrated remarkable capabilities in 2D vision, motivating their extension to 3D scene understanding. Recent studies represent 3D scenes as…
PatchCue: Enhancing Vision-Language Model Reasoning with Patch-Based Visual Cues
Yukun Qi, Pei Fu, Hang Li +5
Vision-Language Models (VLMs) have achieved remarkable progress on a wide range of challenging multimodal understanding and reasoning tasks. However, existing reasoning paradigms,…
Robix: A Unified Model for Robot Interaction, Reasoning and Planning
Huang Fang, Mengxi Zhang, Heng Dong +6
We introduce Robix, a unified model that integrates robot reasoning, task planning, and natural language interaction within a single vision-language architecture. Acting as the hig…
Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
Lin Long, Yichen He, Wentao Ye +5
We introduce M3-Agent, a novel multimodal agent framework equipped with long-term memory. Like humans, M3-Agent can process real-time visual and auditory inputs to build and update…