8 papers · 1 filter
TetherCache: Stabilizing Autoregressive Long-Form Video Generation with Gated Recall and Trusted Alignment
Yu Meng, Xiangyang Luo, Letian Li +5
Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generate…
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
Houcheng Jiang, Jiajun Fu, Junfeng Fang +4
Multimodal large language models are increasingly expected to perform thinking with images, yet existing visual latent reasoning methods still rely on explicit textual chain-of-tho…
HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
Qi Cai, Jingwen Chen, Chengmin Gao +22
The evolution of visual generative models has long been constrained by fragmented architectures relying on disjoint text encoders and external VAEs. In this report, we present HiDr…
Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization
Kaiyuan Li, Xiaoyue Chen, Chen Gao +2
Large Vision-Language Models (LVLMs) have shown impressive performance across multi-modal tasks by encoding images into thousands of tokens. However, the large number of image toke…
RoboScape: Physics-informed Embodied World Model
Yu Shang, Xin Zhang, Yinzhou Tang +4
World models have become indispensable tools for embodied intelligence, serving as powerful simulators capable of generating realistic robotic videos while addressing critical data…
PLPHP: Per-Layer Per-Head Vision Token Pruning for Efficient Large Vision-Language Models
Yu Meng, Kaiyuan Li, Chenran Huang +4
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across a range of multimodal tasks. However, their inference efficiency is constrained by the large n…