7 papers
TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training
Yinyi Luo, Wenwen Wang, Hayes Bai +6
Recent advances in unified multimodal models (UMMs) have led to a proliferation of architectures capable of understanding, generating, and editing across visual and textual modalit…
LatentUMM: Dual Latent Alignment for Unified Multimodal Models
Yinyi Luo, Wenwen Wang, Hayes Bai +2
Unified multimodal models (UMMs) achieve strong performance in both understanding and generation by learning a shared latent space, yet they often exhibit functional inconsistency…
Self-Corrected Image Generation with Explainable Latent Rewards
Yinyi Luo, Hrishikesh Gokhale, Marios Savvides +2
Despite significant progress in text-to-image generation, aligning outputs with complex prompts remains challenging, particularly for fine-grained semantics and spatial relations.…
KnowledgeSmith: Uncovering Knowledge Updating in LLMs with Model Editing and Unlearning
Yinyi Luo, Zhexian Zhou, Hao Chen +4
Knowledge editing and machine unlearning are two popular approaches for large language models (LLMs) to stay up-to-date. However, the knowledge updating mechanism of LLMs remains l…
Image Tokenizer Needs Post-Training
Kai Qiu, Xiang Li, Hao Chen +7
Recent image generative models typically capture the image distribution in a pre-constructed latent space, relying on a frozen image tokenizer. However, there exists a significant…
SOLAR: Scalable Optimization of Large-scale Architecture for Reasoning
Chen Li, Yinyi Luo, Anudeep Bolimera +4
Large Language Models excel in reasoning yet often rely on Chain-of-Thought prompts, limiting performance on tasks demanding more nuanced topological structures. We present SOLAR (…