2 papers
cs.CL2025
Unveiling Reasoning Thresholds in Language Models: Scaling, Fine-Tuning, and Interpretability through Attention Maps
Yen-Che Hsiao, Abhishek Dutta
This study investigates the in-context learning capabilities of various decoder-only transformer-based language models with different model sizes and training data, including GPT2,…
cs.CL2024
Efficient transformer with reinforced position embedding for language models
Yen-Che Hsiao, Abhishek Dutta
In this paper, we propose an efficient transformer architecture that uses reinforced positional embedding to obtain superior performance with half the number of encoder decoder lay…