3 papers
cs.CV2026
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
Junha Song, Byeongho Heo, Geonmo Gu +3
When humans describe a visual scene, they do not process the entire image uniformly; instead, they selectively fixate on regions relevant to their intended description. In contrast…
cs.IR2026
MuCo: Multi-turn Contrastive Learning for Multimodal Embedding Model
Geonmo Gu, Byeongho Heo, Jaemyung Yu +7
Universal Multimodal embedding models built on Multimodal Large Language Models (MLLMs) have traditionally employed contrastive learning, which aligns representations of query-targ…
cs.CV2026
Grounding World Simulation Models in a Real-World Metropolis
Junyoung Seo, Hyunwook Choi, Minkyung Kwon +10
What if a world simulation model could render not an imagined environment but a city that actually exists? Prior generative world models synthesize visually plausible yet artificia…