activity
20222024
most citedVisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

131 citations · 162 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV20244 cited

LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation

Jiaxiang Tang, Zhaoxi Chen, Xiaokang Chen +3

3D content creation has achieved significant progress in terms of both quality and speed. Although current feed-forward models can produce 3D objects in seconds, their resolution i…

cs.CV20239 cited

Interactive Segment Anything NeRF with Feature Imitation

Xiaokang Chen, Jiaxiang Tang, Diwen Wan +2

This paper investigates the potential of enhancing Neural Radiance Fields (NeRF) with semantics to expand their applications. Although NeRF has been proven useful in real-world app…

cs.CV2023131 cited

VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Wenhai Wang, Zhe Chen, Xiaokang Chen +8

Large language models (LLMs) have notably accelerated progress towards artificial general intelligence (AGI), with their impressive zero-shot capacity for user-tailored tasks, endo…

cs.CV2023

Real-time 3D Semantic Scene Completion Via Feature Aggregation and Conditioned Prediction

Xiaokang Chen, Yajie Xing, Gang Zeng

Semantic Scene Completion (SSC) aims to simultaneously predict the volumetric occupancy and semantic category of a 3D scene. In this paper, we propose a real-time semantic scene co…

cs.LG2023

Graph Signal Sampling for Inductive One-Bit Matrix Completion: a Closed-form Solution

Chao Chen, Haoyu Geng, Gang Zeng +4

Inductive one-bit matrix completion is motivated by modern applications such as recommender systems, where new users would appear at test stage with the ratings consisting of only…

cs.CV202218 cited

Conditional DETR V2: Efficient Detection Transformer with Box Queries

Xiaokang Chen, Fangyun Wei, Gang Zeng +1

In this paper, we are interested in Detection Transformer (DETR), an end-to-end object detection approach based on a transformer encoder-decoder architecture without hand-crafted p…