8 papers
ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models
Minho Park, Kinam Kim, Junha Hyung +5
Diffusion and flow matching models have emerged as powerful robot policies, enabling Vision-Language-Action (VLA) models to generalize across diverse scenes and instructions. Yet,…
CA-LoRA: Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation
Minho Park, Sunghyun Park, Jungsoo Lee +5
This paper addresses the challenge of data scarcity in semantic segmentation by generating datasets through text-to-image (T2I) generation models, reducing image acquisition and la…
Not the Example, but the Process: How Self-Generated Examples Enhance LLM Reasoning
Daehoon Gwak, Minseo Jung, Junwoo Park +4
Recent studies have shown that Large Language Models (LLMs) can improve their reasoning performance through self-generated few-shot examples, achieving results comparable to manual…
EgoX: Egocentric Video Generation from a Single Exocentric Video
Taewoong Kang, Kinam Kim, Dohyeon Kim +3
Egocentric perception enables humans to experience and understand the world directly from their own point of view. Translating exocentric (third-person) videos into egocentric (fir…
SphereDiff: Tuning-free 360° Static and Dynamic Panorama Generation via Spherical Latent Representation
Minho Park, Taewoong Kang, Jooyeol Yun +2
The increasing demand for AR/VR applications has highlighted the need for high-quality content, such as 360° live wallpapers. However, generating high-quality 360° panoramic cont…
Reward-Weighted Sampling: Enhancing Non-Autoregressive Characteristics in Masked Diffusion LLMs
Daehoon Gwak, Minseo Jung, Junwoo Park +4
Masked diffusion models (MDMs) offer a promising non-autoregressive alternative for large language modeling. Standard decoding methods for MDMs, such as confidence-based sampling,…