most citedCVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning

1 citations · 1 across the 6 of their papers we have counts for

collaborators

7 papers

cs.RO2026

SG-VLA: Learning Spatially-Grounded Vision-Language-Action Models for Mobile Manipulation

Ruisen Tu, Arth Shukla, Sohyun Yoo +5

Vision-Language-Action (VLA) models show promise for robotic control, yet performance in complex household environments remains sub-optimal. Mobile manipulation requires reasoning…

cs.CV2026

AnyHand: A Large-Scale Synthetic Dataset for RGB(-D) Hand Pose Estimation

Chen Si, Yulin Liu, Bo Ai +4

We present AnyHand, a large-scale synthetic dataset designed to advance the state of the art in 3D hand pose estimation. While recent works with foundation approaches have shown th…

cs.CV2025★ 1 cited

CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning

Zeyuan Chen, Xiang Zhang, Haiyang Xu +2

We present a central-peripheral vision-inspired framework (CVP), a simple yet effective multimodal model for spatial reasoning that draws inspiration from the two types of human vi…

cs.CV2025

VideoNSA: Native Sparse Attention Scales Video Understanding

Enxin Song, Wenhao Chai, Shusheng Yang +5

Video understanding in multimodal language models remains limited by context length: models often miss key transition frames and struggle to maintain coherence across long time sca…

cs.CV2025

OverLayBench: A Benchmark for Layout-to-Image Generation with Dense Overlaps

Bingnan Li, Chen-Yu Wang, Haiyang Xu +7

Despite steady progress in layout-to-image generation, current methods still struggle with layouts containing significant overlap between bounding boxes. We identify two primary ch…

cs.CV2025

DepR: Depth Guided Single-view Scene Reconstruction with Instance-level Diffusion

Qingcheng Zhao, Xiang Zhang, Haiyang Xu +4

We propose DepR, a depth-guided single-view scene reconstruction framework that integrates instance-level diffusion within a compositional paradigm. Instead of reconstructing the e…