activity
20182026
most citedBeyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

565 citations · 853 across the 78 of their papers we have counts for

collaborators
Showing 2018 · cs.CLShow all

7 papers · 2 filters

cs.CL2018

Towards End-to-end Automatic Code-Switching Speech Recognition

Genta Indra Winata, Andrea Madotto, Chien-Sheng Wu +1

Speech recognition in mixed language has difficulties to adapt end-to-end framework due to the lack of data and overlapping phone sets, for example in words such as "one" in Englis…

cs.CL2018

Learn to Code-Switch: Data Augmentation using Copy Mechanism on Language Modeling

Genta Indra Winata, Andrea Madotto, Chien-Sheng Wu +1

Building large-scale datasets for training code-switching language models is challenging and very expensive. To alleviate this problem using parallel corpus has been a major workar…

cs.CL2018

Learning Comment Generation by Leveraging User-Generated Data

Zhaojiang Lin, Genta Indra Winata, Pascale Fung

Existing models on open-domain comment generation are difficult to train, and they produce repetitive and uninteresting responses. The problem is due to multiple and contradictory…

cs.CL2018

Handling Imbalanced Dataset in Multi-label Text Categorization using Bagging and Adaptive Boosting

Genta Indra Winata, Masayu Leylia Khodra

Imbalanced dataset is occurred due to uneven distribution of data available in the real world such as disposition of complaints on government offices in Bandung. Consequently, mult…

cs.CL2018

Attention-Based LSTM for Psychological Stress Detection from Spoken Language Using Distant Supervision

Genta Indra Winata, Onno Pepijn Kampman, Pascale Fung

We propose a Long Short-Term Memory (LSTM) with attention mechanism to classify psychological stress from self-conducted interview transcriptions. We apply distant supervision by a…

cs.CL2018

Code-Switching Language Modeling using Syntax-Aware Multi-Task Learning

Genta Indra Winata, Andrea Madotto, Chien-Sheng Wu +1

Lack of text data has been the major issue on code-switching language modeling. In this paper, we introduce multi-task learning based language model which shares syntax representat…