1 citations · 4 across the 17 of their papers we have counts for
5 papers · 1 filter
Factorized Visual Tokenization and Generation
Zechen Bai, Jianxiong Gao, Ziteng Gao +4
Visual tokenizers are fundamental to image generation. They convert visual data into discrete tokens, enabling transformer-based models to excel at image generation. Despite their…
VideoSAM: Open-World Video Segmentation
Pinxue Guo, Zixu Zhao, Jianxiong Gao +5
Video segmentation is essential for advancing robotics and autonomous driving, particularly in open-world settings where continuous perception and object association across video f…
MinD-3D++: Advancing fMRI-Based 3D Reconstruction with High-Quality Textured Mesh Generation and a Comprehensive Dataset
Jianxiong Gao, Yanwei Fu, Yuqian Fu +3
Reconstructing 3D visuals from functional Magnetic Resonance Imaging (fMRI) data, introduced as Recon3DMind, is of significant interest to both cognitive neuroscience and computer…
LAC-Net: Linear-Fusion Attention-Guided Convolutional Network for Accurate Robotic Grasping Under the Occlusion
Jinyu Zhang, Yongchong Gu, Jianxiong Gao +5
This paper addresses the challenge of perceiving complete object shapes through visual perception. While prior studies have demonstrated encouraging outcomes in segmenting the visi…
Hyper-Transformer for Amodal Completion
Jianxiong Gao, Xuelin Qian, Longfei Liang +2
Amodal object completion is a complex task that involves predicting the invisible parts of an object based on visible segments and background information. Learning shape priors is…