1 citations · 1 across the 15 of their papers we have counts for
14 papers · 1 filter
Fine-Grained Multi Image Object Hallucination Benchmark
Joonki Min, Chaeyun Kim, Hyungwook Choi +4
Multimodal Large Language Models (MLLMs) are increasingly deployed in multi-image scenarios requiring complex reasoning across visual contexts. However, current MLLMs remain fundam…
Domain Generalization via Text-Anchored Information Bottleneck
Eunyi Lyou, Yunjeong Choi, Junho Lee +1
Visual recognition models often fail when deployed in new environments. Domain Generalization (DG) addresses this by learning representations that remain invariant to environment-s…
Equivariant Latent Alignment via Flow Matching under Group Symmetries
Sunghyun Kim, Jaehoon Hahm, Jeongwoo Shin +1
Geometry-aware generative models and novel view synthesis approaches have shown strong potential in visual fidelity and consistency. In parallel, equivariant representation learnin…
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
Geo Ahn, Jiwook Han, Youngrae Kim +2
Fine-tuning MLLMs for Video Temporal Grounding (VTG) often improves in-domain performance but degrades sharply under domain shift. In this work, we find that this failure is primar…
Geometry-Aware Image Flow Matching
Junho Lee, Kwanseok Kim, Joonseok Lee
Recent advances in generative models highlight the power of geometry-aware modeling in manifold-constrained settings. Yet, for natural images, the field remains confined to Euclide…
ArtSplat: Feed-Forward Articulated 3D Gaussian Splatting from Sparse Multi-State Uncalibrated Views
Inseo Lee, Yoonji Kim, Eugene Sohn +4
Articulated object reconstruction from sparse-view images is an ill-posed problem that requires simultaneous inference of geometry and underlying articulation structure. Existing m…