activity
20242026
collaborators

8 papers

cs.CV2026

Low-Resolution Editing is All You Need for High-Resolution Editing

Junsung Lee, Hyunsoo Lee, Yong Jae Lee +1

High-resolution content creation is rapidly emerging as a central challenge in both the vision and graphics communities. Images serve as the most fundamental modality for visual ex…

cs.CV2026

Relational Visual Similarity

Thao Nguyen, Sicheng Mo, Krishna Kumar Singh +6

Humans do not just see attribute similarity -- we also see relational similarity. An apple is like a peach because both are reddish fruit, but the Earth is also like a peach: its c…

cs.CV2025

Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration

Sicheng Mo, Thao Nguyen, Richard Zhang +7

In this work, we explore an untapped signal in diffusion model inference. While all previous methods generate images independently at inference, we instead ask if samples can be ge…

cs.CV2025

YoChameleon: Personalized Vision and Language Generation

Thao Nguyen, Krishna Kumar Singh, Jing Shi +3

Large Multimodal Models (e.g., GPT-4, Gemini, Chameleon) have evolved into powerful tools with millions of users. However, they remain generic models and lack personalized knowledg…

cs.CV2025

X-Fusion: Introducing New Modality to Frozen Large Language Models

Sicheng Mo, Thao Nguyen, Xun Huang +9

We propose X-Fusion, a framework that extends pretrained Large Language Models (LLMs) for multimodal tasks while preserving their language capabilities. X-Fusion employs a dual-tow…

cs.CV2025

Efficient LLaMA-3.2-Vision by Trimming Cross-attended Visual Features

Jewon Lee, Ki-Ung Song, Seungmin Yang +6

Visual token reduction lowers inference costs caused by extensive image features in large vision-language models (LVLMs). Unlike relevant studies that prune tokens in self-attentio…