18 papers
Global Geometry Is Not Enough for Vision Representations
Jiwan Chung, Seon Joo Kim
A common assumption in representation learning is that globally well-distributed embeddings support robust and generalizable representations. This focus has shaped both training ob…
Rethinking State Tracking in Recurrent Models Through Error Control Dynamics
Jiwan Chung, Heechan Choi, Seon Joo Kim
The theory of state tracking in recurrent architectures has predominantly focused on expressive capacity: whether a fixed architecture can theoretically realize a set of symbolic t…
What MLLMs Learn about When they Learn about Multimodal Reasoning
Jiwan Chung, Neel Joshi, Pratyusha Sharma +2
Evaluation of multimodal reasoning models is typically reduced to a single accuracy score, implicitly treating reasoning as a unitary capability. We introduce MathLens, a benchmark…
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
Jiwan Chung, Junhyeok Kim, Siyeol Kim +3
When thinking with images, humans rarely rely on a single glance: they revisit visual evidence while reasoning. In contrast, most Multimodal Language Models encode an image once to…
Teaching Metric Distance to Discrete Autoregressive Language Models
Jiwan Chung, Saejin Kim, Yongrae Jo +3
Large language models (LLMs) operate as autoregressive predictors over discrete token vocabularies, a formulation that has enabled their adaptation far beyond natural language to v…
GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance
Junhyeok Kim, Jaewoo Park, Junhee Park +5
For people affected by blindness and low vision (BLV), safe and independent navigation remains a major challenge, impacting over 2.2 billion individuals worldwide. Although multimo…