Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Discrete JEPA: Learning Discrete Token Representations without Reconstruction
Junyeob Baek, Hosung Lee, Christopher Hoang +2
The cornerstone of cognitive intelligence lies in extracting hidden patterns from observations and leveraging these principles to systematically predict future outcomes. However, c…
cs.CV2025
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
Donghwan Chi, Hyomin Kim, Yoonjin Oh +7
Recently, multimodal large language models (MLLMs) have emerged as a key approach in achieving artificial general intelligence. In particular, vision-language MLLMs have been devel…