1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.CV2025
MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
Yuqi Pang, Bowen Yang, Yun Cao +3
Vision large language models (VLLMs) are focusing primarily on handling complex and fine-grained visual information by incorporating advanced vision encoders and scaling up visual…
cs.AI2025★ 1 cited
Ming-Omni: A Unified Multimodal Model for Perception and Generation
Inclusion AI, Biao Gong, Cheng Zou +55
We propose Ming-Omni, a unified multimodal model capable of processing images, text, audio, and video, while demonstrating strong proficiency in both speech and image generation. M…
cs.SD2022
A Policy-based Approach to the SpecAugment Method for Low Resource E2E ASR
Rui Li, Guodong Ma, Dexin Zhao +3
SpecAugment is a very effective data augmentation method for both HMM and E2E-based automatic speech recognition (ASR) systems. Especially, it also works in low-resource scenarios.…