4 papers · 1 filter
Towards Compositional Generalization of LLMs via Skill Taxonomy Guided Data Synthesis
Yifan Wei, Li Du, Xiaoyan Yu +2
Large Language Models (LLMs) and agent-based systems often struggle with compositional generalization due to a data bottleneck in which complex skill combinations follow a long-tai…
MoCE: Adaptive Mixture of Contextualization Experts for Byte-based Neural Machine Translation
Langlin Huang, Mengyu Bu, Yang Feng
Byte-based machine translation systems have shown significant potential in massively multilingual settings. Unicode encoding, which maps each character to specific byte(s), elimina…
Integrating Multi-scale Contextualized Information for Byte-based Neural Machine Translation
Langlin Huang, Yang Feng
Subword tokenization is a common method for vocabulary building in Neural Machine Translation (NMT) models. However, increasingly complex tasks have revealed its disadvantages. Fir…
TA&AT: Enhancing Task-Oriented Dialog with Turn-Level Auxiliary Tasks and Action-Tree Based Scheduled Sampling
Longxiang Liu, Xiuxing Li, Yang Feng
Task-oriented dialog systems have witnessed substantial progress due to conversational pre-training techniques. Yet, two significant challenges persist. First, most systems primari…