3 papers
cs.LG2026
Learning to Theorize the World from Observation
Doojin Baek, Gyubin Lee, Junyeob Baek +2
What does it mean to understand the world? Contemporary world models often operationalize understanding as accurate future prediction in latent or observation space. Developmental…
cs.CV2026
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
Donghwan Chi, Hyomin Kim, Yoonjin Oh +7
Recently, multimodal large language models (MLLMs) have emerged as a key approach in achieving artificial general intelligence. In particular, vision-language MLLMs have been devel…
cs.CV2025
Discrete JEPA: Learning Discrete Token Representations without Reconstruction
Junyeob Baek, Hosung Lee, Christopher Hoang +2
The cornerstone of cognitive intelligence lies in extracting hidden patterns from observations and leveraging these principles to systematically predict future outcomes. However, c…