6 papers · 1 filter
COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models
Ziqi Zhou, Weize Quan, Mining Tan +6
Composition is a high-level visual intent that governs where subjects are placed and how a scene is organized, yet current unified multimodal models remain unreliable at fine-grain…
The Wittgensteinian Representation Hypothesis: Is Language the Attractor of Multimodal Convergence?
Zhaoyang Zhang, Run Shao, Dongyue Wu +4
Understanding why independently trained neural networks from different modalities converge toward shared representations, and where this convergence leads, remains an open question…
GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks
Cong Chen, Kaixiang Ji, Hao Zhong +9
Autonomous agents for long-sequence Graphical User Interface tasks are hindered by sparse rewards and the intractable credit assignment problem. To address these challenges, we int…
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
Inclusion AI, :, Fudong Wang +12
Recent advancements in Multimodal Large Language Models (MLLMs), particularly through Reinforcement Learning with Verifiable Rewards (RLVR), have significantly enhanced their reaso…
Ming-Omni: A Unified Multimodal Model for Perception and Generation
Inclusion AI, Biao Gong, Cheng Zou +55
We propose Ming-Omni, a unified multimodal model capable of processing images, text, audio, and video, while demonstrating strong proficiency in both speech and image generation. M…
Large Multimodal Model Compression via Efficient Pruning and Distillation at AntGroup
Maolin Wang, Yao Zhao, Jiajia Liu +5
The deployment of Large Multimodal Models (LMMs) within AntGroup has significantly advanced multimodal tasks in payment, security, and advertising, notably enhancing advertisement…