activity
20182021
most citedEfficient Domain Adaptation of Language Models via Adaptive Tokenization

1 citations · 1 across the 1 of their papers we have counts for

collaborators

5 papers

cs.CL20211 cited

Efficient Domain Adaptation of Language Models via Adaptive Tokenization

Vin Sachidananda, Jason S. Kessler, Yi-an Lai

Contextual embedding-based language models trained on large data sets, such as BERT and RoBERTa, provide strong performance across a wide range of tasks and are ubiquitous in moder…

cs.CL2021

Regression Bugs Are In Your Model! Measuring, Reducing and Analyzing Regressions In NLP Model Updates

Yuqing Xie, Yi-an Lai, Yuanjun Xiong +2

Behavior of deep neural networks can be inconsistent between different versions. Regressions during model update are a common cause of concern that often over-weigh the benefits in…

cs.CL2020

Diversity, Density, and Homogeneity: Quantitative Characteristic Metrics for Text Collections

Yi-An Lai, Xuan Zhu, Yi Zhang +1

Summarizing data samples by quantitative measures has a long history, with descriptive statistics being a case in point. However, as natural language processing methods flourish, t…

cs.CL2019

Goal-Embedded Dual Hierarchical Model for Task-Oriented Dialogue Generation

Yi-An Lai, Arshit Gupta, Yi Zhang

Hierarchical neural networks are often used to model inherent structures within dialogues. For goal-oriented dialogues, these models miss a mechanism adhering to the goals and negl…

cs.IR2018

Attribute-aware Collaborative Filtering: Survey and Classification

Wen-Hao Chen, Chin-Chi Hsu, Yi-An Lai +3

Attribute-aware CF models aims at rating prediction given not only the historical rating from users to items, but also the information associated with users (e.g. age), items (e.g.…