1 citations · 1 across the 1 of their papers we have counts for
1 paper
Anna C. Marbut, John W. Chandler, Travis J. Wheeler
It is generally thought that transformer-based large language models benefit from pre-training by learning generic linguistic knowledge that can be focused on a specific task durin…