14 papers · 1 filter
Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models
Yitong Chen, Shiduo Zhang, Jingjing Gong +1
Generating diverse images from sparse text is hard; generating compact actions from rich observations is easier. From the condition-target view, Vision-Language-Action (VLA) thus a…
IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder
Yitong Chen, Zijie Diao, Junke Wang +5
Built on pretrained vision foundation models (VFMs), representation autoencoders (RAEs) have recently emerged as a promising approach for constructing semantically rich latent spac…
Channel-wise Vector Quantization
Wei Song, Tianhang Wang, Yitong Chen +5
We present Channel-wise Vector Quantization (CVQ), a novel image tokenization paradigm that replaces patch-wise tokens with channel-wise tokens. Unlike conventional vector quantiza…
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization
Zhuohan Liu, Wujian Peng, Yitong Chen +1
Despite the rapid progress of text-to-image (T2I) models, generating images that accurately reflect complex compositional prompts (covering attribute bindings, object relationships…
DecQ: Detail-Condensing Queries for Enhanced Reconstruction and Generation in Representation Autoencoders
Tianhang Wang, Yitong Chen, Wei Song +3
Representation Autoencoders (RAEs) leverage frozen vision foundation models (VFMs) as tokenizer encoders, providing robust high-level representations that facilitate fast convergen…
DPC-VQA: Decoupling Quality Perception and Residual Calibration for Video Quality Assessment
Xinyue Li, Shubo Xu, Zhichao Zhang +3
Recent multimodal large language models (MLLMs) have shown promising performance on video quality assessment (VQA) tasks. However, adapting them to new scenarios remains expensive…