5 papers
Retrieve What's Missing: Coverage-Maximizing Retrieval for Consistent Long Video Generation
Minseok Joo, Dogyun Park, Taehoon Lee +2
Maintaining long-term geometric consistency remains challenging for long-horizon autoregressive video generation. Memory-augmented generative models address this by retrieving hist…
Mitigating State Aliasing in Vision-Language-Action Models via Inverse Dynamics Learning
Kyujin Lee, Injae Kim, Jihwan Park +3
Vision-Language-Action (VLA) models have emerged as a promising framework that unifies perception, reasoning, and control for robot manipulation by adapting pretrained vision-langu…
F4Splat: Feed-Forward Predictive Densification for Feed-Forward 3D Gaussian Splatting
Injae Kim, Chaehyeon Kim, Minseong Bae +2
Feed-forward 3D Gaussian Splatting methods enable single-pass reconstruction and real-time rendering. However, they typically adopt rigid pixel-to-Gaussian or voxel-to-Gaussian pip…
Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
Dogyun Park, Taehoon Lee, Minseok Joo +1
Recently, Flow Matching models have pushed the boundaries of high-fidelity data generation across a wide range of domains. It typically employs a single large network to learn the…
Generative Subgraph Retrieval for Knowledge Graph-Grounded Dialog Generation
Jinyoung Park, Minseok Joo, Joo-Kyung Kim +1
Knowledge graph-grounded dialog generation requires retrieving a dialog-relevant subgraph from the given knowledge base graph and integrating it with the dialog history. Previous w…