6 papers
Token Warping Helps MLLMs Look from Nearby Viewpoints
Phillip Y. Lee, Chanho Park, Mingue Park +3
Can warping tokens, rather than pixels, help multimodal large language models (MLLMs) understand how a scene appears from a nearby viewpoint? While MLLMs perform well on visual rea…
BoxSplitGen: A Generative Model for 3D Part Bounding Boxes in Varying Granularity
Juil Koo, Wei-Tung Lin, Chanho Park +2
Human creativity follows a perceptual process, moving from abstract ideas to finer details during creation. While 3D generative models have advanced dramatically, models specifical…
DiffusionRollout: Uncertainty-Aware Rollout Planning in Long-Horizon PDE Solving
Seungwoo Yoo, Juil Koo, Daehyeon Choi +1
We propose DiffusionRollout, a novel selective rollout planning strategy for autoregressive diffusion models, aimed at mitigating error accumulation in long-horizon predictions of…
Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
Juil Koo, Daehyeon Choi, Sangwoo Youn +2
Vision Language Models (VLMs) excel at visual question answering (VQA) but remain limited to snapshot vision, reasoning from static images. In contrast, embodied agents require amb…
BézierFlow: Learning Bézier Stochastic Interpolant Schedulers for Few-Step Generation
Yunhong Min, Juil Koo, Seungwoo Yoo +1
We introduce BézierFlow, a lightweight training approach for few-step generation with pretrained diffusion and flow models. BézierFlow achieves a 2-3x performance improvement for s…
VideoHandles: Editing 3D Object Compositions in Videos Using Video Generative Priors
Juil Koo, Paul Guerrero, Chun-Hao Paul Huang +2
Generative methods for image and video editing use generative models as priors to perform edits despite incomplete information, such as changing the composition of 3D objects shown…