activity
20232026
most citedLanguage Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch

13 citations · 28 across the 26 of their papers we have counts for

collaborators
Showing 2023 · cs.CLShow all

7 papers · 2 filters

cs.CL2023

Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use

Yuhan Chen, Ang Lv, Ting-En Lin +5

In this paper, we demonstrate that an inherent waveform pattern in the attention allocation of large language models (LLMs) significantly affects their performance in tasks demandi…

cs.CL2023★ 13 cited

Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch

Le Yu, Bowen Yu, Haiyang Yu +2

In this paper, we unveil that Language Models (LMs) can acquire new capabilities by assimilating parameters from homologous models without retraining or GPUs. We first introduce DA…

cs.CL2023★ 1 cited

Diversify Question Generation with Retrieval-Augmented Style Transfer

Qi Gou, Zehua Xia, Bowen Yu +4

Given a textual passage and an answer, humans are able to ask questions with various expressions, but this ability is still challenging for most question generation (QG) systems. E…

cs.CL2023★ 2 cited

Improving Question Generation with Multi-level Content Planning

Zehua Xia, Qi Gou, Bowen Yu +4

This paper addresses the problem of generating questions from a given context and an answer, specifically focusing on questions that require multi-hop reasoning across an extended…

cs.CL2023★ 1 cited

Exploring Large Language Models for Multi-Modal Out-of-Distribution Detection

Yi Dai, Hao Lang, Kaisheng Zeng +2

Out-of-distribution (OOD) detection is essential for reliable and trustworthy machine learning. Recent multi-modal OOD detection leverages textual information from in-distribution…

cs.CL2023

Constructive Large Language Models Alignment with Diverse Feedback

Tianshu Yu, Ting-En Lin, Yuchuan Wu +3

In recent research on large language models (LLMs), there has been a growing emphasis on aligning these models with human values to reduce the impact of harmful content. However, c…