activity
20242026
collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI2026

COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models

Ziqi Zhou, Weize Quan, Mining Tan +6

Composition is a high-level visual intent that governs where subjects are placed and how a scene is organized, yet current unified multimodal models remain unreliable at fine-grain…

cs.AI2026

The Wittgensteinian Representation Hypothesis: Is Language the Attractor of Multimodal Convergence?

Zhaoyang Zhang, Run Shao, Dongyue Wu +4

Understanding why independently trained neural networks from different modalities converge toward shared representations, and where this convergence leads, remains an open question…

cs.AI2025

GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks

Cong Chen, Kaixiang Ji, Hao Zhong +9

Autonomous agents for long-sequence Graphical User Interface tasks are hindered by sparse rewards and the intractable credit assignment problem. To address these challenges, we int…

cs.AI2025

M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning

Inclusion AI, :, Fudong Wang +12

Recent advancements in Multimodal Large Language Models (MLLMs), particularly through Reinforcement Learning with Verifiable Rewards (RLVR), have significantly enhanced their reaso…

cs.AI2025

Ming-Omni: A Unified Multimodal Model for Perception and Generation

Inclusion AI, Biao Gong, Cheng Zou +55

We propose Ming-Omni, a unified multimodal model capable of processing images, text, audio, and video, while demonstrating strong proficiency in both speech and image generation. M…

cs.AI2024

Large Multimodal Model Compression via Efficient Pruning and Distillation at AntGroup

Maolin Wang, Yao Zhao, Jiajia Liu +5

The deployment of Large Multimodal Models (LMMs) within AntGroup has significantly advanced multimodal tasks in payment, security, and advertising, notably enhancing advertisement…