6 papers
On the Convergence of Gradient Descent on Learning Transformers with Residual Connections
Zhen Qin, Jinxin Zhou, Jiachen Jiang +1
Transformer models have emerged as fundamental tools across various scientific and engineering disciplines, owing to their outstanding performance in diverse applications. Despite…
Improving Visual Discriminability of CLIP for Training-Free Open-Vocabulary Semantic Segmentation
Jinxin Zhou, Jiachen Jiang, Zhihui Zhu
Extending CLIP models to semantic segmentation remains challenging due to the misalignment between their image-level pre-training objectives and the pixel-level visual understandin…
From Compression to Expression: A Layerwise Analysis of In-Context Learning
Jiachen Jiang, Yuxin Dong, Jinxin Zhou +1
In-context learning (ICL) enables large language models (LLMs) to adapt to new tasks without weight updates by learning from demonstration sequences. While ICL shows strong empiric…
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
Jiachen Jiang, Jinxin Zhou, Bo Peng +2
Achieving better alignment between vision embeddings and Large Language Models (LLMs) is crucial for enhancing the abilities of Multimodal LLMs (MLLMs), particularly for recent mod…
Tracing Representation Progression: Analyzing and Enhancing Layer-Wise Similarity
Jiachen Jiang, Jinxin Zhou, Zhihui Zhu
Analyzing the similarity of internal representations has been an important technique for understanding the behavior of deep neural networks. Most existing methods for analyzing the…
Cat-AIR: Content and Task-Aware All-in-One Image Restoration
Jiachen Jiang, Tianyu Ding, Ke Zhang +5
All-in-one image restoration seeks to recover high-quality images from various types of degradation using a single model, without prior knowledge of the corruption source. However,…