212 citations · 448 across the 46 of their papers we have counts for
9 papers · 1 filter
M-VAR: Decoupled Scale-wise Autoregressive Modeling for High-Quality Image Generation
Sucheng Ren, Yaodong Yu, Nataniel Ruiz +3
There exists recent work in computer vision, named VAR, that proposes a new autoregressive paradigm for image generation. Diverging from the vanilla next-token prediction, VAR stru…
Adventurer: Optimizing Vision Mamba Architecture Designs for Efficiency
Feng Wang, Timing Yang, Yaodong Yu +7
In this work, we introduce the Adventurer series models where we treat images as sequences of patch tokens and employ uni-directional language models to learn visual representation…
From Pixels to Objects: A Hierarchical Approach for Part and Object Segmentation Using Local and Global Aggregation
Yunfei Xie, Cihang Xie, Alan Yuille +1
In this paper, we introduce a hierarchical transformer-based model designed for sophisticated image segmentation tasks, effectively bridging the granularity of part segmentation wi…
MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine
Yunfei Xie, Ce Zhou, Lang Gao +8
This paper introduces MedTrinity-25M, a comprehensive, large-scale multimodal dataset for medicine, covering over 25 million images across 10 modalities with multigranular annotati…
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning
Sucheng Ren, Hongru Zhu, Chen Wei +3
This paper presents a new self-supervised video representation learning framework, ARVideo, which autoregressively predicts the next video token in a tailored sequence order. Two k…
Mamba-R: Vision Mamba ALSO Needs Registers
Feng Wang, Jiahao Wang, Sucheng Ren +6
Similar to Vision Transformers, this paper identifies artifacts also present within the feature maps of Vision Mamba. These artifacts, corresponding to high-norm tokens emerging in…