5 papers
LLaViDA: A Large Language Vision Driving Assistant for Explicit Reasoning and Enhanced Trajectory Planning
Yudong Liu, Spencer Hallyburton, Jiwoo Kim +8
Trajectory planning is a fundamental yet challenging component of autonomous driving. End-to-end planners frequently falter under adverse weather, unpredictable human behavior, or…
Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
Dohun Lee, Hyeonho Jeong, Jiwook Kim +2
Video diffusion models have advanced rapidly in the recent years as a result of series of architectural innovations (e.g., diffusion transformers) and use of novel training objecti…
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
Seonho Lee, Jiho Choi, Inha Kang +3
Vision-Language Models (VLMs) have shown remarkable performance on diverse visual and linguistic tasks, yet they remain fundamentally limited in their understanding of 3D spatial s…
Doppler Correspondence: Non-Iterative Scan Matching With Doppler Velocity-Based Correspondence
Jiwoo Kim, Geunsik Bae, Changseung Kim +3
Achieving successful scan matching is essential for LiDAR odometry. However, in challenging environments with adverse weather conditions or repetitive geometric patterns, LiDAR odo…
Scribble-Guided Diffusion for Training-free Text-to-Image Generation
Seonho Lee, Jiho Choi, Seohyun Lim +2
Recent advancements in text-to-image diffusion models have demonstrated remarkable success, yet they often struggle to fully capture the user's intent. Existing approaches using te…