5 papers
Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
Yi Yang, Xueqi Li, Yiyang Chen +7
Recent advances in Vision-Language-Action (VLA) models demonstrate that visual signals can effectively complement sparse action supervisions. However, letting VLA directly predict…
LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation
Ethan Chern, Zhulin Hu, Bohao Tang +4
Real-time video generation via diffusion is essential for building general-purpose multimodal interactive AI systems. However, the simultaneous denoising of all video frames with b…
Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
Bohao Tang, Yan Ma, Fei Zhang +6
Chart understanding presents a critical test to the reasoning capabilities of Vision-Language Models (VLMs). Prior approaches face critical limitations: some rely on external tools…
LIMO: Less is More for Reasoning
Yixin Ye, Zhen Huang, Yang Xiao +3
We challenge the prevailing assumption that complex reasoning in large language models (LLMs) necessitates massive training data. We demonstrate that sophisticated mathematical rea…
Thinking with Generated Images
Ethan Chern, Zhulin Hu, Steffi Chern +5
We present Thinking with Generated Images, a novel paradigm that fundamentally transforms how large multimodal models (LMMs) engage with visual reasoning by enabling them to native…