15 papers
ReWorld: Learning Better Representations for World Action Models
Tianze Xia, Lijun Zhou, Kaixin Xiong +9
World Action Models (WAMs) model future environment evolution under action conditioning, offering a scalable paradigm for autonomous driving. However, existing approaches focus lar…
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
Kangsheng Duan, Ziyang Xu, Wenyu Liu +3
While 10B-level industrial foundation models have pushed the boundaries of image inpainting, their prohibitive computational costs severely hinder practical deployment. Constructin…
Food-R1: A Unified Multi-Task Food Vision-Language Model with Reinforcement Learning
Yu Zhu, Yongkang Li, Wenjie Zhu +5
Recent studies have explored Vision-Language Models (VLMs) for food analysis. However, most existing methods rely primarily on supervised fine-tuning (SFT), which often limits reas…
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
Lunbin Zeng, Jingfeng Yao, Bencheng Liao +3
Diffusion-based decoding has recently emerged as an appealing alternative to autoregressive (AR) generation, offering the potential to update multiple tokens in parallel and reduce…
Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation
Tianheng Cheng, Xinggang Wang, Junchao Liao +1
Semantic segmentation is a fundamental problem in computer vision and it requires high-resolution feature maps for dense prediction. Current coordinate-guided low-resolution featur…
4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
Xianfeng Wu, Yajing Bai, Minghan Li +5
Constructing 4D language fields is crucial for embodied AI, augmented/virtual reality, and 4D scene understanding, as they provide enriched semantic representations of dynamic envi…