2 papers
cs.CV2026
Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
Muyang Li, Yucheng Liu, Jianbo Ma +3
Vision-Language Models (VLMs) have enhanced traditional LLMs with visual capabilities through the integration of vision encoders. While recent works have explored various combinati…
cs.MM2025
Gotta Hear Them All: Towards Sound Source Aware Audio Generation
Wei Guo, Heng Wang, Jianbo Ma +1
Audio synthesis has broad applications in multimedia. Recent advancements have made it possible to generate relevant audios from inputs describing an audio scene, such as images or…