13 papers
When 2D Tasks Meet 1D Serialization: On Serialization Friction in Structured Tasks
Chung-Hsiang Lo, Lu Li, Diji Yang +4
In the LLM era, many symbolic and structured problems are presented to models through 1D text serialization. Yet some such problems are natively two-dimensional: their relevant rel…
GRIT: Teaching MLLMs to Think with Images
Yue Fan, Xuehai He, Diji Yang +6
Recent studies have demonstrated the efficacy of using Reinforcement Learning (RL) in building reasoning models that articulate chains of thoughts prior to producing final answers.…
Grid2Matrix: Revealing Digital Agnosia in Vision-Language Models
Yunkai Zhang, Linda Li, Yingxin Cui +5
Vision-Language Models (VLMs) excel on many multimodal reasoning benchmarks, but these evaluations often do not require an exhaustive readout of the image and can therefore obscure…
Retrieval-Augmented Robots via Retrieve-Reason-Act
Izat Temiraliev, Diji Yang, Yi Zhang
To achieve general-purpose utility, we argue that robots must evolve from passive executors into active Information Retrieval users. In strictly zero-shot settings where no prior d…
Classroom Final Exam: An Instructor-Tested Reasoning Benchmark
Chongyang Gao, Diji Yang, Shuyan Zhou +4
We introduce CFE-Bench (Classroom Final Exam), a multimodal benchmark for evaluating the reasoning capabilities of large language models across more than 20 STEM domains. CFE-Bench…
Shared Nature, Unique Nurture: PRISM for Pluralistic Reasoning via In-context Structure Modeling
Guancheng Tu, Shiyang Zhang, Tianyu Zhang +2
Large Language Models (LLMs) are converging towards a singular Artificial Hivemind, where shared Nature (pre-training priors) result in a profound collapse of distributional divers…