1 citations · 1 across the 4 of their papers we have counts for
6 papers
Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning
Sanghyuk Chun, Olga Russakovsky
Multimodal learning has seen remarkable progress, particularly with large-scale pre-training across various modalities. Most current approaches are built on the assumption of a det…
CoMet: Context and Multiplicity Decomposition for Multimodal Uncertainty Estimation
Sanghyuk Chun, William Yang, Amaya Dharmasiri +1
Uncertainty estimation has been a long-standing challenge in AI models; it amounts to "knowing what you don't know," and metacognition is notoriously difficult even for humans (cf.…
Mitigating Cross-Image Information Leakage in Multi-Image Understanding with Large Vision-Language Models
Yeji Park, Minyoung Lee, Sanghyuk Chun +1
Large Vision-Language Models (LVLMs) exhibit strong performance on single-image tasks. However, their performance degrades significantly when handling multi-image inputs. While thi…
Reinforced Fast Weights with Next-Sequence Prediction
Hee Seung Hwang, Xindi Wu, Sanghyuk Chun +1
Fast weight architectures offer a promising alternative to attention-based transformers for long-context modeling by maintaining constant memory overhead regardless of context leng…
Seeing What You Say: Expressive Image Generation from Speech
Jiyoung Lee, Song Park, Sanghyuk Chun +1
This paper proposes VoxStudio, the first unified and end-to-end speech-to-image model that generates expressive images directly from spoken descriptions by jointly aligning linguis…
Read, Watch and Scream! Sound Generation from Text and Video
Yujin Jeong, Yunji Kim, Sanghyuk Chun +1
Despite the impressive progress of multimodal generative models, video-to-audio generation still suffers from limited performance and limits the flexibility to prioritize sound syn…