5 papers
Introspective Diffusion Language Models
Yifan Yu, Yuqing Jian, Junxiong Wang +12
Diffusion language models promise parallel generation, yet still lag behind autoregressive (AR) models in quality. We stem this gap to a failure of introspective consistency: AR mo…
PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models
Chak-Wing Mak, Guanyu Zhu, Boyi Zhang +16
Modern foundational Multimodal Large Language Models (MLLMs) and video world models have advanced significantly in mathematical, common-sense, and visual reasoning, but their grasp…
ARM2: Adaptive Reasoning Model with Vision Understanding and Executable Code
Jian Xie, Zhendong Chu, Aoxiao Zhong +5
Large Reasoning Models (LRMs) often suffer from the ``over-thinking'' problem, generating unnecessarily long reasoning on simple tasks. Some strategies have been proposed to mitiga…
IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
Yang Chen, Minghao Liu, Yufan Shen +18
The webpage-to-code task requires models to understand visual representations of webpages and generate corresponding code. However, existing benchmarks primarily focus on static sc…
Empower Vision Applications with LoRA LMM
Liang Mi, Weijun Wang, Wenming Tu +10
Large Multimodal Models (LMMs) have shown significant progress in various complex vision tasks with the solid linguistic and reasoning capacity inherited from large language models…