4 papers
UniCon-Former: Unified Convolution Transformer is All You Need for Hand Gesture Recognition
Mallika Garg, Debashis Ghosh, Pyari Mohan Pradhan
Convolutional Neural Networks (CNNs) capture local features efficiently but struggle with global context due to their limited receptive field. On the other hand, transformers effec…
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
Inclusion AI, :, Fudong Wang +12
Recent advancements in Multimodal Large Language Models (MLLMs), particularly through Reinforcement Learning with Verifiable Rewards (RLVR), have significantly enhanced their reaso…
Ming-Omni: A Unified Multimodal Model for Perception and Generation
Inclusion AI, Biao Gong, Cheng Zou +55
We propose Ming-Omni, a unified multimodal model capable of processing images, text, audio, and video, while demonstrating strong proficiency in both speech and image generation. M…
M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive Performance
Qingpei Guo, Kaiyou Song, Zipeng Feng +9
We present M2-omni, a cutting-edge, open-source omni-MLLM that achieves competitive performance to GPT-4o. M2-omni employs a unified multimodal sequence modeling framework, which e…