activity
20242026
most citedMultiplicity is an Inevitable and Inherent Challenge in Multimodal Learning

1 citations · 1 across the 4 of their papers we have counts for

collaborators

6 papers

cs.LG20261 cited

Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning

Sanghyuk Chun, Olga Russakovsky

Multimodal learning has seen remarkable progress, particularly with large-scale pre-training across various modalities. Most current approaches are built on the assumption of a det…

cs.LG2026

CoMet: Context and Multiplicity Decomposition for Multimodal Uncertainty Estimation

Sanghyuk Chun, William Yang, Amaya Dharmasiri +1

Uncertainty estimation has been a long-standing challenge in AI models; it amounts to "knowing what you don't know," and metacognition is notoriously difficult even for humans (cf.…

cs.CV2026

Mitigating Cross-Image Information Leakage in Multi-Image Understanding with Large Vision-Language Models

Yeji Park, Minyoung Lee, Sanghyuk Chun +1

Large Vision-Language Models (LVLMs) exhibit strong performance on single-image tasks. However, their performance degrades significantly when handling multi-image inputs. While thi…

cs.CL2026

Reinforced Fast Weights with Next-Sequence Prediction

Hee Seung Hwang, Xindi Wu, Sanghyuk Chun +1

Fast weight architectures offer a promising alternative to attention-based transformers for long-context modeling by maintaining constant memory overhead regardless of context leng…

eess.AS2025

Seeing What You Say: Expressive Image Generation from Speech

Jiyoung Lee, Song Park, Sanghyuk Chun +1

This paper proposes VoxStudio, the first unified and end-to-end speech-to-image model that generates expressive images directly from spoken descriptions by jointly aligning linguis…

cs.CV2024

Read, Watch and Scream! Sound Generation from Text and Video

Yujin Jeong, Yunji Kim, Sanghyuk Chun +1

Despite the impressive progress of multimodal generative models, video-to-audio generation still suffers from limited performance and limits the flexibility to prioritize sound syn…