works on

From the 1 of 24 linked papers with an AI index.

activity
20242026
collaborators

24 papers

cs.CV2026

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning

Jiahao Shao, Yuanbo Yang, Yiyi Liao +3

Tool-augmented vision-language models increasingly "think with images": they call crop, zoom, or code tools and reason over the returned pixels. However, recent work using blind te…

cs.RO2026

Native Video-Action Pretraining for Generalizable Robot Control

Qihang Zhang, Lin Li, Luyao Zhang +26

The paper introduces LingBot-VA 2.0, a video-action foundation model designed specifically for robot control, featuring a semantic visual-action tokenizer, causal pretraining, a sp…

cs.CV2026

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Shuailei Ma, Jiaqi Liao, Xinyang Wang +24

Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inheren…

cs.CV2026

Infinite Worlds with Versatile Interactions

Zelin Gao, Qiuyu Wang, Jiapeng Zhu +17

We present LingBot-World 2.0 (also known as LingBot-World-Infinity), an advanced iteration of LingBot-World featuring four distinct upgrades. (1) Our model achieves an unbounded in…

cs.CV2026

Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator

Zihan Wang, Seungjun Lee, Yinghao Xu +1

Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destinations reliably in the real world. However, progress remains co…

cs.CV2026

Vision Pretraining for Dense Spatial Perception

Zelin Fu, Bin Tan, Changjiang Sun +6

Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observat…