6 papers
HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
HyperAI Team, Yuchen Liu, Kaiyang Han +26
Current multimodal large lanauge models possess strong perceptual and reasoning capabilities, however high computational and memory requirements make them difficult to deploy direc…
EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence
Ding Zou, Feifan Wang, Mengyu Ge +17
The realization of Artificial General Intelligence (AGI) necessitates Embodied AI agents capable of robust spatial perception, effective task planning, and adaptive execution in ph…
Igniting VLMs toward the Embodied Space
Andy Zhai, Brae Liu, Bruno Fang +17
While foundation models show remarkable progress in language and vision, existing vision-language models (VLMs) still have limited spatial and embodiment understanding. Transferrin…
Kwai Keye-VL 1.5 Technical Report
Biao Yang, Bin Wen, Boyang Ding +58
In recent years, the development of Large Language Models (LLMs) has significantly advanced, extending their capabilities to multimodal tasks through Multimodal Large Language Mode…
VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction
Hao Wang, Eiki Murata, Lingfang Zhang +11
Recent advances in multimodal large language models (MLLMs) have significantly enhanced video understanding capabilities, opening new possibilities for practical applications. Yet…
GenVP: Generating Visual Puzzles with Contrastive Hierarchical VAEs
Kalliopi Basioti, Pritish Sahu, Qingze Tony Liu +3
Raven's Progressive Matrices (RPMs) is an established benchmark to examine the ability to perform high-level abstract visual reasoning (AVR). Despite the current success of algorit…