collaborators

6 papers

eess.AS2026

StepAudio 2.5 Technical Report

Bin Lin, Bo Zhao, Boyong Wu +98

Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks.…

cs.CV2026

Vision Foundation Models as Generalist Tokenizers for Image Generation

Anlin Zheng, Qi Han, Xin Wen +5

In this work, we explore the largely unexplored direction of building a generalist image tokenizer directly on top of a frozen vision foundation model (VFM). To build this tokenize…

cs.CV2026

Dropping Anchor and Spherical Harmonics for Sparse-view Gaussian Splatting

Shuangkang Fang, I-Chao Shen, Xuanyang Zhang +5

Recent 3D Gaussian Splatting (3DGS) Dropout methods address overfitting under sparse-view conditions by randomly nullifying Gaussian opacities. However, we identify a neighbor comp…

cs.CV2025

IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction

Hao Li, Zhengyu Zou, Fangfu Liu +8

Humans naturally perceive the geometric structure and semantic content of a 3D world as intertwined dimensions, enabling coherent and accurate understanding of complex scenes. Howe…

cs.CV2025

Vision Foundation Models as Effective Visual Tokenizers for Autoregressive Image Generation

Anlin Zheng, Xin Wen, Xuanyang Zhang +5

In this work, we present a novel direction to build an image tokenizer directly on top of a frozen vision foundation model, which is a largely underexplored area. Specifically, we…

cs.CV2025

NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale

NextStep Team, Chunrui Han, Guopeng Li +47

Prevailing autoregressive (AR) models for text-to-image generation either rely on heavy, computationally-intensive diffusion models to process continuous image tokens, or employ ve…