Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
Inclusion AI, :, Bowen Ma +73
We propose Ming-Flash-Omni, an upgraded version of Ming-Omni, built upon a sparser Mixture-of-Experts (MoE) variant of Ling-Flash-2.0 with 100 billion total parameters, of which on…
cs.CV2025
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
Jisheng Dang, Huilin Song, Junbin Xiao +6
Grounded Video Question Answering (Grounded VideoQA) requires aligning textual answers with explicit visual evidence. However, modern multimodal models often rely on linguistic pri…