From the 1 of 7 linked papers with an AI index.
7 papers
Understanding Knowledge Transfer Mechanism in Heterogeneous MLLM Fusion: A Simple Linear Approach
Yinghao Hou, Jiahe Fan, Yuanhao Pu +2
The paper introduces a simple linear probe called Cross-Scale Directional Parameter Injection (CDPI) to study how knowledge transfers when heterogeneous multimodal large language m…
Rethinking Heterogeneous LLM Merging: A Weighted Model Averaging Perspective
Jiahe Fan, Yinghao Hou, Si Chen +3
Can large language models with substantially different parameter spaces be merged by direct weighted averaging, without training or semantic alignment? Existing heterogeneous fusio…
Optimization-Free Topological Sort for Causal Discovery via the Schur Complement of Score Jacobians
Rui Wu, Hong Xie
Continuous causal discovery typically couples representation learning with structural optimization via non-convex acyclicity penalties, which subjects solvers to local optima and r…
Model Specific Task Similarity for Vision Language Model Selection via Layer Conductance
Wei Yang, Hong Xie, Tao Tan +3
While open sourced Vision-Language Models (VLMs) have proliferated, selecting the optimal pretrained model for a specific downstream task remains challenging. Exhaustive evaluation…
Demystifying Design Choices of Reinforcement Fine-tuning: A Batched Contextual Bandit Learning Perspective
Hong Xie, Xiao Hu, Tao Tan +5
The reinforcement fine-tuning area is undergoing an explosion papers largely on optimizing design choices. Though performance gains are often claimed, inconsistent conclusions also…
Rethinking Reinforcement fine-tuning of LLMs: A Multi-armed Bandit Learning Perspective
Xiao Hu, Hong Xie, Tao Tan +2
A large number of heuristics have been proposed to optimize the reinforcement fine-tuning of LLMs. However, inconsistent claims are made from time to time, making this area elusive…