collaborators

5 papers

cs.RO2026

SERF: Spatiotemporal Environment and Robot Feature Map for Long-Horizon Mobile Manipulation

Sunghwan Kim, Byeonghyun Pak, Kehan Long +2

Long-horizon robot mobile manipulation requires continual reasoning about localization, environment changes, and task progress, all of which are challenging to infer from image obs…

cs.CV2026

Rethinking FID Through the Geometry of the Reference Dataset

Yunghee Lee, Byeonghyun Pak

Fréchet Inception Distance (FID) is widely used to evaluate image generators, yet lower FID does not always correspond to better sample quality. We show that this mismatch depends…

cs.CV2026

Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding

Byeongju Woo, Zilin Wang, Byeonghyun Pak +2

Vision-language models such as CLIP often struggle to faithfully understand long, detail-rich captions, relying on dominant scene cues while overlooking fine-grained visual evidenc…

cs.CV2026

Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition

Seokmin Lee, Yunghee Lee, Byeonghyun Pak +1

For robotic agents operating in dynamic environments, learning visual state representations from streaming video observations is essential for sequential decision making. Recent se…

cs.CV2025

Tortoise and Hare Guidance: Accelerating Diffusion Model Inference with Multirate Integration

Yunghee Lee, Byeonghyun Pak, Junwha Hong +1

In this paper, we propose Tortoise and Hare Guidance (THG), a training-free strategy that accelerates diffusion sampling while maintaining high-fidelity generation. We demonstrate…