activity
20242026
collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2026

ScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generation

Jiawei Zhang, Hongsong Wang, Pan Zhou

Text-driven 3D indoor scene generation has advanced from dataset-bound layout modeling to open-vocabulary synthesis with large language and vision-language models. Yet existing met…

cs.CV2026

PathRelax: Parallel-Path Relaxed Speculative Jacobi Decoding for Accelerating Auto-Regressive Text-to-Image Generation

Haodong Lei, Hongsong Wang, Bingxuan Dai +1

The growing need for high-resolution image generation in autoregressive text-to-image models has resulted in extended token sequences, significantly increasing computational costs…

cs.CV2026

Bilingual Text-to-Motion Generation: A New Benchmark and Baselines

Wanjiang Weng, Xiaofeng Tan, Xiangbo Shu +3

Text-to-motion generation holds significant potential for cross-linguistic applications, yet it is hindered by the lack of bilingual datasets and the poor cross-lingual semantic un…

cs.CV2025

Fast Inference of Visual Autoregressive Model with Adjacency-Adaptive Dynamical Draft Trees

Haodong Lei, Hongsong Wang, Xin Geng +2

Autoregressive (AR) image models achieve diffusion-level quality but suffer from sequential inference, requiring approximately 2,000 steps for a 576x576 image. Speculative decoding…

cs.CV2025

ReAlign: Text-to-Motion Generation via Step-Aware Reward-Guided Alignment

Wanjiang Weng, Xiaofeng Tan, Junbo Wang +3

Text-to-motion generation, which synthesizes 3D human motions from text inputs, holds immense potential for applications in gaming, film, and robotics. Recently, diffusion-based me…

cs.CV2025

Dragging with Geometry: From Pixels to Geometry-Guided Image Editing

Xinyu Pu, Hongsong Wang, Jie Gui +1

Interactive point-based image editing serves as a controllable editor, enabling precise and flexible manipulation of image content. However, most drag-based methods operate primari…