18 papers
Beyond Routing Saturation: A Long-Horizon Class-Incremental Perspective on Expert Routing in Multimodal Continual Instruction Tuning
Huiyu Yi, Yongqi Xu, Bogang Zhang +5
Multimodal Continual Instruction Tuning (MCIT) enables multimodal large language models to acquire new tasks sequentially while retaining previously learned capabilities. Many rece…
Redirecting the Flow: Image Customization through Attention Distribution Shift
Jie Li, Suorong Yang, Jian Zhao +1
Subject-driven image customization aims to generate images that not only follow textual instructions but also preserve the identity of a given reference subject. Existing approache…
A Geometric Measure of Linear Separability for Neural Representations
Yi Wei, Xuan Qi, Furao Shen
Modern neural classifiers commonly rely on linear readouts, yet predictive metrics alone do not characterize the class-wise geometry of the representations on which such readouts o…
Beyond What to Select: A Plug-and-play Oscillatory Data-Volume Scheduling for Efficient Model Training
Suorong Yang, Hanqi Zhu, Hai Gan +4
Data selection accelerates training by identifying representative training data while preserving model performance. However, existing methods mainly focus on designing sample-impor…
Data Agent: Learning to Select Data via End-to-End Dynamic Optimization
Suorong Yang, Fangjian Su, Hai Gan +5
Dynamic Data selection aims to accelerate training by prioritizing informative samples during online training. However, existing methods typically rely on task-specific handcrafted…
AffineLens: Capturing the Continuous Piecewise Affine Functions of Neural Networks
Yi Wei, Xuan Qi, Furao Shen +3
Piecewise affine neural networks (PANNs) provide a principled geometric perspective on neural network expressivity by characterizing the input--output map as a continuous piecewise…