activity
20072023
most citedIndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP

32 citations · 187 across the 35 of their papers we have counts for

collaborators
Showing 2017Show all

5 papers · 1 filter

cs.CL201710 cited

Capturing Long-range Contextual Dependencies with Memory-enhanced Conditional Random Fields

Fei Liu, Timothy Baldwin, Trevor Cohn

Despite successful applications across a broad range of NLP tasks, conditional random fields ("CRFs"), in particular the linear-chain variant, are only able to model local features…

cs.CL2017

Continuous Representation of Location for Geolocation and Lexical Dialectology using Mixture Density Networks

Afshin Rahimi, Timothy Baldwin, Trevor Cohn

We propose a method for embedding two-dimensional locations in a continuous vector space using a neural network-based model incorporating mixtures of Gaussian distributions, presen…

cs.CL2017

An Automatic Approach for Document-level Topic Model Evaluation

Shraey Bhatia, Jey Han Lau, Timothy Baldwin

Topic models jointly learn topics and document-level topic distribution. Extrinsic evaluation of topic models tends to focus exclusively on topic-level evaluation, e.g. by assessin…

cs.CL2017

Topically Driven Neural Language Model

Jey Han Lau, Timothy Baldwin, Trevor Cohn

Language models are typically applied at the sentence level, without access to the broader document context. We present a neural language model that incorporates document context i…

cs.CL2017

Context-Aware Prediction of Derivational Word-forms

Ekaterina Vylomova, Ryan Cotterell, Timothy Baldwin +1

Derivational morphology is a fundamental and complex characteristic of language. In this paper we propose the new task of predicting the derivational form of a given base-form lemm…