4 papers
XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments
Kangan Qian, ChuChu Xie, Yang Zhong +13
Vision-Language-Action (VLA) models drive next-generation autonomous systems, but training them requires scalable, high-quality annotations from complex environments. Current cloud…
Q-DeepSight: Incentivizing Thinking with Images for Image Quality Assessment and Refinement
Xudong Li, Jiaxi Tan, Ziyin Zhou +6
Image Quality Assessment (IQA) models are increasingly deployed as perceptual critics to guide generative models and image restoration. This role demands not only accurate scores b…
BADiff: Bandwidth Adaptive Diffusion Model
Xi Zhang, Hanwei Zhu, Yan Zhong +2
In this work, we propose a novel framework to enable diffusion models to adapt their generation quality based on real-time network bandwidth constraints. Traditional diffusion mode…
OmniQuality-R: Advancing Reward Models Through All-Encompassing Quality Assessment
Yiting Lu, Fengbin Guan, Yixin Gao +8
Current visual evaluation approaches are typically constrained to a single task. To address this, we propose OmniQuality-R, a unified reward modeling framework that transforms mult…