5 papers
TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
Shu-Hao Zhang, Wei-Cheng Tang, Chen Wu +5
Recent years have witnessed an increasing interest in image-text contrastive modeling, exemplified by models such as Contrastive Language-Image Pretraining (CLIP). In this paper, w…
IL3D: A Large-Scale Indoor Layout Dataset for LLM-Driven 3D Scene Generation
Wenxu Zhou, Kaixuan Nie, Hang Du +5
In this study, we present IL3D, a large-scale dataset meticulously designed for large language model (LLM)-driven 3D scene generation, addressing the pressing demand for diverse, h…
Beyond Instance Consistency: Investigating View Diversity in Self-supervised Learning
Huaiyuan Qin, Muli Yang, Siyuan Hu +4
Self-supervised learning (SSL) conventionally relies on the instance consistency paradigm, assuming that different views of the same image can be treated as positive pairs. However…
Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots
Minghuan Liu, Zhengbang Zhu, Xiaoshen Han +12
Modern robotic manipulation primarily relies on visual observations in a 2D color space for skill learning but suffers from poor generalization. In contrast, humans, living in a 3D…
IQPFR: An Image Quality Prior for Blind Face Restoration and Beyond
Peng Hu, Chunming He, Lei Xu +5
Blind Face Restoration (BFR) addresses the challenge of reconstructing degraded low-quality (LQ) facial images into high-quality (HQ) outputs. Conventional approaches predominantly…