activity
20232026
most citedVideoSAM: Open-World Video Segmentation

1 citations · 4 across the 17 of their papers we have counts for

collaborators
Showing 2024Show all

5 papers · 1 filter

cs.CV2024

Factorized Visual Tokenization and Generation

Zechen Bai, Jianxiong Gao, Ziteng Gao +4

Visual tokenizers are fundamental to image generation. They convert visual data into discrete tokens, enabling transformer-based models to excel at image generation. Despite their…

cs.CV2024★ 1 cited

VideoSAM: Open-World Video Segmentation

Pinxue Guo, Zixu Zhao, Jianxiong Gao +5

Video segmentation is essential for advancing robotics and autonomous driving, particularly in open-world settings where continuous perception and object association across video f…

cs.CV2024

MinD-3D++: Advancing fMRI-Based 3D Reconstruction with High-Quality Textured Mesh Generation and a Comprehensive Dataset

Jianxiong Gao, Yanwei Fu, Yuqian Fu +3

Reconstructing 3D visuals from functional Magnetic Resonance Imaging (fMRI) data, introduced as Recon3DMind, is of significant interest to both cognitive neuroscience and computer…

cs.RO2024★ 1 cited

LAC-Net: Linear-Fusion Attention-Guided Convolutional Network for Accurate Robotic Grasping Under the Occlusion

Jinyu Zhang, Yongchong Gu, Jianxiong Gao +5

This paper addresses the challenge of perceiving complete object shapes through visual perception. While prior studies have demonstrated encouraging outcomes in segmenting the visi…

cs.CV2024

Hyper-Transformer for Amodal Completion

Jianxiong Gao, Xuelin Qian, Longfei Liang +2

Amodal object completion is a complex task that involves predicting the invisible parts of an object based on visible segments and background information. Learning shape priors is…