4 papers
Efficient and High-Quality Depth Estimation via Pixel-Space Diffusion with Linear Attention
Bingde Liu, Wu Ran, Jinglei Zhang +2
This work presents , a inear-ttention-based xel-pace generative framework that achieves efficient and high-fidelity…
MOSS-VL Technical Report
Pengyu Wang, Chenkun Tan, Shaojun Zhou +29
We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across th…
Guiding a Diffusion Model by Swapping Its Tokens
Weijia Zhang, Yuehao Liu, Shanyan Guan +4
Classifier-Free Guidance (CFG) is a widely used inference-time technique to boost the image quality of diffusion models. Yet, its reliance on text conditions prevents its use in un…
Cross-Architecture Distillation Made Simple with Redundancy Suppression
Weijia Zhang, Yuehao Liu, Wu Ran +1
We describe a simple method for cross-architecture knowledge distillation, where the knowledge transfer is cast into a redundant information suppression formulation. Existing metho…