adapters 1large language models 1mixture of experts 1multi-task learning 1parameter-efficient fine-tuning 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CL2026
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
Congmin Zheng, Jiachen Zhu, Jianghao Lin +6
Process Reward Models (PRMs) play a central role in evaluating and guiding multi-step reasoning in large language models (LLMs), especially for mathematical problem solving. Howeve…
cs.CL2024
Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Liangchen Luo, Yinxiao Liu, Rosanne Liu +9
Complex multi-step reasoning tasks, such as solving mathematical problems or generating code, remain a significant hurdle for even the most advanced large language models (LLMs). V…
cs.CL2024
MoDE: Effective Multi-task Parameter Efficient Fine-Tuning with a Mixture of Dyadic Experts
Lin Ning, Harsh Lara, Meiqi Guo +1
The paper introduces MoDE, a multi-task parameter-efficient fine-tuning method for large language models that shares down-projection matrices and uses rank-one adapters with task r…