4 papers
From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion
Cheng Chen, Yuyu Guo, Pengpeng Zeng +4
Vision-Language Models (VLMs) create a severe visual feature bottleneck by using a crude, asymmetric connection that links only the output of the vision encoder to the input of the…
Reversible Inversion for Training-Free Exemplar-guided Image Editing
Yuke Li, Lianli Gao, Ji Zhang +5
Exemplar-guided Image Editing (EIE) aims to modify a source image according to a visual reference. Existing approaches often require large-scale pre-training to learn relationships…
Sim-and-Human Co-training for Data-Efficient and Generalizable Robotic Manipulation
Kaipeng Fang, Weiqing Liang, Yuyang Li +5
Synthetic simulation data and real-world human data provide scalable alternatives to circumvent the prohibitive costs of robot data collection. However, these sources suffer from t…
Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models
Chenhang Cui, Gelei Deng, An Zhang +5
Recent advances in Large Vision-Language Models (LVLMs) have showcased strong reasoning abilities across multiple modalities, achieving significant breakthroughs in various real-wo…