17 papers · 1 filter
Thinking Ahead: Foresight Intelligence in MLLMs and World Model
Zhantao Gong, Liaoyuan Fan, Qing Guo +3
In this work, we define Foresight Intelligence as the capability to anticipate and interpret future events-an ability essential for applications such as autonomous driving, yet lar…
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images
Qirui Wang, Jingyi He, Yining Pan +3
Spatial reasoning (SR), the ability to infer 3D spatial information from 2D inputs, is essential for real-world applications such as embodied AI and autonomous driving. However, ex…
Grounding by Remembering: Cross-Scene and In-Scene Memory for 3D Functional Affordances
Qirui Wang, Jingyi He, Yining Pan +2
Functional affordance grounding requires more than recognizing an object: an agent must localize the specific region that supports an interaction, such as the handle to pull or the…
Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
Yongyi Su, Haojie Zhang, Shijie Li +11
Multimodal large language models (MLLMs) have advanced rapidly in recent years. However, existing approaches for vision tasks often rely on indirect representations, such as genera…
DiffPCN: Latent Diffusion Model Based on Multi-view Depth Images for Point Cloud Completion
Zijun Li, Hongyu Yan, Shijie Li +4
Latent diffusion models (LDMs) have demonstrated remarkable generative capabilities across various low-level vision tasks. However, their potential for point cloud completion remai…
AD-FM: Multimodal LLMs for Anomaly Detection via Multi-Stage Reasoning and Fine-Grained Reward Optimization
Jingyi Liao, Yongyi Su, Rong-Cheng Tu +6
While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities across diverse domains, their application to specialized anomaly detection (AD) remains constrain…