1 citations · 1 across the 1 of their papers we have counts for
2 papers
cs.CV2026
From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation
Haowen Gu, Gensheng Pei, Junzhu Mao +3
Although Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in Medical Visual Question Answering (Med-VQA), their reliance on global image features o…
cs.AI2024★ 1 cited
Yuan 2.0-M32: Mixture of Experts with Attention Router
Shaohua Wu, Jiangang Luo, Xi Chen +12
Yuan 2.0-M32, with a similar base architecture as Yuan-2.0 2B, uses a mixture-of-experts architecture with 32 experts of which 2 experts are active. A new router network, Attention…