1 citations · 1 across the 6 of their papers we have counts for
7 papers
SG-VLA: Learning Spatially-Grounded Vision-Language-Action Models for Mobile Manipulation
Ruisen Tu, Arth Shukla, Sohyun Yoo +5
Vision-Language-Action (VLA) models show promise for robotic control, yet performance in complex household environments remains sub-optimal. Mobile manipulation requires reasoning…
AnyHand: A Large-Scale Synthetic Dataset for RGB(-D) Hand Pose Estimation
Chen Si, Yulin Liu, Bo Ai +4
We present AnyHand, a large-scale synthetic dataset designed to advance the state of the art in 3D hand pose estimation. While recent works with foundation approaches have shown th…
CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning
Zeyuan Chen, Xiang Zhang, Haiyang Xu +2
We present a central-peripheral vision-inspired framework (CVP), a simple yet effective multimodal model for spatial reasoning that draws inspiration from the two types of human vi…
VideoNSA: Native Sparse Attention Scales Video Understanding
Enxin Song, Wenhao Chai, Shusheng Yang +5
Video understanding in multimodal language models remains limited by context length: models often miss key transition frames and struggle to maintain coherence across long time sca…
OverLayBench: A Benchmark for Layout-to-Image Generation with Dense Overlaps
Bingnan Li, Chen-Yu Wang, Haiyang Xu +7
Despite steady progress in layout-to-image generation, current methods still struggle with layouts containing significant overlap between bounding boxes. We identify two primary ch…
DepR: Depth Guided Single-view Scene Reconstruction with Instance-level Diffusion
Qingcheng Zhao, Xiang Zhang, Haiyang Xu +4
We propose DepR, a depth-guided single-view scene reconstruction framework that integrates instance-level diffusion within a compositional paradigm. Instead of reconstructing the e…