collaborators

7 papers

cs.AI2026

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training

Yinyi Luo, Wenwen Wang, Hayes Bai +6

Recent advances in unified multimodal models (UMMs) have led to a proliferation of architectures capable of understanding, generating, and editing across visual and textual modalit…

cs.CV2026

LatentUMM: Dual Latent Alignment for Unified Multimodal Models

Yinyi Luo, Wenwen Wang, Hayes Bai +2

Unified multimodal models (UMMs) achieve strong performance in both understanding and generation by learning a shared latent space, yet they often exhibit functional inconsistency…

cs.MM2026

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning

Hayes Bai, Yinyi Luo, Wenwen Wang +2

Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture. However, it remains underexplored how to effectively coordinate these t…

cs.CL2026

Can Vision Replace Text in Working Memory? Evidence from Spatial n-Back in Vision-Language Models

Sichu Liang, Hongyu Zhu, Wenwen Wang +1

Working memory is a central component of intelligent behavior, providing a dynamic workspace for maintaining and updating task-relevant information. Recent work has used n-back tas…

cs.CV2025

Evading Data Provenance in Deep Neural Networks

Hongyu Zhu, Sichu Liang, Wenwen Wang +3

Modern over-parameterized deep models are highly data-dependent, with large scale general-purpose and domain-specific datasets serving as the bedrock for rapid advancements. Howeve…

cs.CV2025

Revisiting Data Auditing in Large Vision-Language Models

Hongyu Zhu, Sichu Liang, Wenwen Wang +5

With the surge of large language models (LLMs), Large Vision-Language Models (VLMs)--which integrate vision encoders with LLMs for accurate visual grounding--have shown great poten…