5 papers
VLANeXt: Recipes for Building Strong VLA Models
Xiao-Ming Wu, Bin Fan, Kang Liao +6
Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding from Vision-Language Models for gen…
ConceptSeg-R1: Segment Any Concept via Meta-Reinforcement Learning
Yuan Zhao, Youwei Pang, Jiaming Zuo +10
Recent progress in promptable segmentation has shifted visual perception from object-level localization toward concept-level understanding. However, the notion of a concept remains…
Learning Spatial Decay for Vision Transformers
Yuxin Mao, Zhen Qin, Jinxing Zhou +4
Vision Transformers (ViTs) have revolutionized computer vision, yet their self-attention mechanism lacks explicit spatial inductive biases, leading to suboptimal performance on spa…
Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective
Yuxin Mao, Zhen Qin, Jinxing Zhou +6
Autoregressive (AR) models have garnered significant attention in image generation for their ability to effectively capture both local and global structures within visual data. How…
Instance-Level Moving Object Segmentation from a Single Image with Events
Zhexiong Wan, Bin Fan, Le Hui +2
Moving object segmentation plays a crucial role in understanding dynamic scenes involving multiple moving objects, while the difficulties lie in taking into account both spatial te…