collaborators

11 papers

cs.AI2026

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning

Wencheng Ye, Yi Bin, Yujuan Ding +7

Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it enters the language stack, weakening eviden…

cs.RO2026

WSA: a 3D-Centric World-Spatial-Action Model for Generalizable Robot Control

Jiahao Jiang, Jianing Zhang, Zhenhan Yin +8

Recent advances in embodied AI have established robot foundation models (RFMs) as the dominant approach for generalist robotic systems to date. By leveraging imitation learning on…

cs.CV2026

TIMI: Training-Free Image-to-3D Multi-Instance Generation with Spatial Fidelity

Xiao Cai, Pengpeng Zeng, Ji Zhang +3

Precise spatial fidelity in Image-to-3D multi-instance generation is critical for downstream real-world applications. Recent work attempts to address this by fine-tuning pre-traine…

cs.RO2026

DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models

Siyuan Xu, Tianshi Wang, Fengling Li +2

Vision-Language-Action models (VLAs) have demonstrated strong potential for embodied AI, yet their deployment on resource-limited robots remains challenging due to high memory and…

cs.CV2026

DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset

Hengyu Shen, Tiancheng Gu, Bin Qin +10

Vision-Language Pre-training (VLP) models have achieved remarkable success by leveraging large-scale image-text pairs. While English-centric models like CLIP and SigLIP benefit fro…

cs.RO2026

Non-Markovian Long-Horizon Robot Manipulation via Keyframe Chaining

Yipeng Chen, Wentao Tan, Lei Zhu +4

Existing Vision-Language-Action (VLA) models often struggle to generalize to long-horizon tasks due to their heavy reliance on immediate observations. While recent studies incorpor…