activity
20242026
collaborators
Showing cs.CVShow all

14 papers · 1 filter

cs.CV2026

Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models

Yitong Chen, Shiduo Zhang, Jingjing Gong +1

Generating diverse images from sparse text is hard; generating compact actions from rich observations is easier. From the condition-target view, Vision-Language-Action (VLA) thus a…

cs.CV2026

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder

Yitong Chen, Zijie Diao, Junke Wang +5

Built on pretrained vision foundation models (VFMs), representation autoencoders (RAEs) have recently emerged as a promising approach for constructing semantically rich latent spac…

cs.CV2026

Channel-wise Vector Quantization

Wei Song, Tianhang Wang, Yitong Chen +5

We present Channel-wise Vector Quantization (CVQ), a novel image tokenization paradigm that replaces patch-wise tokens with channel-wise tokens. Unlike conventional vector quantiza…

cs.CV2026

Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization

Zhuohan Liu, Wujian Peng, Yitong Chen +1

Despite the rapid progress of text-to-image (T2I) models, generating images that accurately reflect complex compositional prompts (covering attribute bindings, object relationships…

cs.CV2026

DecQ: Detail-Condensing Queries for Enhanced Reconstruction and Generation in Representation Autoencoders

Tianhang Wang, Yitong Chen, Wei Song +3

Representation Autoencoders (RAEs) leverage frozen vision foundation models (VFMs) as tokenizer encoders, providing robust high-level representations that facilitate fast convergen…

cs.CV2026

DPC-VQA: Decoupling Quality Perception and Residual Calibration for Video Quality Assessment

Xinyue Li, Shubo Xu, Zhichao Zhang +3

Recent multimodal large language models (MLLMs) have shown promising performance on video quality assessment (VQA) tasks. However, adapting them to new scenarios remains expensive…