activity
20232026
collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Rethinking Weight Tying: Pseudo-Inverse Tying for LM Stable Training and Updates

Jian Gu, Aldeida Aleti, Chunyang Chen +1

Weight tying is widely used in compact language models to reduce parameters by sharing the token table between the input embedding and the output projection. However, parameter sha…

cs.CL2025

Beyond Neural Incompatibility: Cross-Scale Knowledge Transfer in Language Models through Latent Semantic Alignment

Jian Gu, Aldeida Aleti, Chunyang Chen +1

Language Models (LMs) encode substantial knowledge in their parameters, yet it remains unclear how to transfer such knowledge in a fine-grained manner, namely parametric knowledge…

cs.CL2025

SeMe: Training-Free Language Model Merging via Semantic Alignment

Jian Gu, Aldeida Aleti, Chunyang Chen +1

Despite the remarkable capabilities of Language Models (LMs) across diverse tasks, no single model consistently outperforms others, necessitating efficient methods to combine their…

cs.CL2024

A Semantic-Aware Layer-Freezing Approach to Computation-Efficient Fine-Tuning of Language Models

Jian Gu, Aldeida Aleti, Chunyang Chen +1

Finetuning language models (LMs) is crucial for adapting the models to downstream data and tasks. However, full finetuning is usually costly. Existing work, such as parameter-effic…

cs.CL2024

Vocabulary-Defined Semantics: Latent Space Clustering for Improving In-Context Learning

Jian Gu, Aldeida Aleti, Chunyang Chen +1

In-context learning enables language models (LM) to adapt to downstream data or tasks by incorporating few samples as demonstrations within the prompts. It offers strong performanc…