3 citations · 3 across the 7 of their papers we have counts for
4 papers · 1 filter
TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training
Yinyi Luo, Wenwen Wang, Hayes Bai +6
Recent advances in unified multimodal models (UMMs) have led to a proliferation of architectures capable of understanding, generating, and editing across visual and textual modalit…
LatentUMM: Dual Latent Alignment for Unified Multimodal Models
Yinyi Luo, Wenwen Wang, Hayes Bai +2
Unified multimodal models (UMMs) achieve strong performance in both understanding and generation by learning a shared latent space, yet they often exhibit functional inconsistency…
UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning
Hayes Bai, Yinyi Luo, Wenwen Wang +2
Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture. However, it remains underexplored how to effectively coordinate these t…
Can Vision Replace Text in Working Memory? Evidence from Spatial n-Back in Vision-Language Models
Sichu Liang, Hongyu Zhu, Wenwen Wang +1
Working memory is a central component of intelligent behavior, providing a dynamic workspace for maintaining and updating task-relevant information. Recent work has used n-back tas…