From the 1 of 8 linked papers with an AI index.
8 papers
Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models
Yitong Chen, Shiduo Zhang, Jingjing Gong +1
The paper proposes a one-step action generation method for vision‑language‑action models, using high‑noise training and a flow‑matching loss, and demonstrates strong performance on…
IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder
Yitong Chen, Zijie Diao, Junke Wang +5
Built on pretrained vision foundation models (VFMs), representation autoencoders (RAEs) have recently emerged as a promising approach for constructing semantically rich latent spac…
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization
Zhuohan Liu, Wujian Peng, Yitong Chen +1
Despite the rapid progress of text-to-image (T2I) models, generating images that accurately reflect complex compositional prompts (covering attribute bindings, object relationships…
DecQ: Detail-Condensing Queries for Enhanced Reconstruction and Generation in Representation Autoencoders
Tianhang Wang, Yitong Chen, Wei Song +3
Representation Autoencoders (RAEs) leverage frozen vision foundation models (VFMs) as tokenizer encoders, providing robust high-level representations that facilitate fast convergen…
INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
Wujian Peng, Lingchen Meng, Yitong Chen +7
Large Multimodal Models (LMMs) have made significant breakthroughs with the advancement of instruction tuning. However, while existing models can understand images and videos at a…
In-depth Analysis on Caching and Pre-fetching in Mixture of Experts Offloading
Shuning Lin, Yifan He, Yitong Chen
In today's landscape, Mixture of Experts (MoE) is a crucial architecture that has been used by many of the most advanced models. One of the major challenges of MoE models is that t…