collaborators

5 papers

cs.CV2026

SJD-PV: Speculative Jacobi Decoding with Phrase Verification for Autoregressive Image Generation

Zhehao Yu, Baoquan Zhang, Bingqi Shan +5

Autoregressive (AR) image models have recently demonstrated remarkable generative capability, but their sequential nature results in significant inference latency. Existing trainin…

cs.CV2025

ST-Booster: An Iterative SpatioTemporal Perception Booster for Vision-and-Language Navigation in Continuous Environments

Lu Yue, Dongliang Zhou, Liang Xie +2

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to navigate unknown, continuous spaces based on natural language instructions. Compared to discre…

cs.SD2025

AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals

Dongliang Zhou, Yakun Zhang, Jinghan Wu +3

The global aging population faces considerable challenges, particularly in communication, due to the prevalence of hearing and speech impairments. To address these, we introduce th…

cs.CV2025

PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation

Sen Wang, Dongliang Zhou, Liang Xie +3

Vision-and-language navigation (VLN) tasks require agents to navigate three-dimensional environments guided by natural language instructions, offering substantial potential for div…

cs.CV2025

LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition

Bowen Hao, Dongliang Zhou, Xiaojie Li +4

Visual speech recognition (VSR), commonly known as lip reading, has garnered significant attention due to its wide-ranging practical applications. The advent of deep learning techn…