32 citations · 176 across the 74 of their papers we have counts for
10 papers · 1 filter
Liputan6: A Large-scale Indonesian Dataset for Text Summarization
Fajri Koto, Jey Han Lau, Timothy Baldwin
In this paper, we introduce a large-scale Indonesian summarization dataset. We harvest articles from Liputan6.com, an online news portal, and obtain 215,827 document-summary pairs.…
IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP
Fajri Koto, Afshin Rahimi, Jey Han Lau +1
Although the Indonesian language is spoken by almost 200 million people and the 10th most spoken language in the world, it is under-represented in NLP research. Previous work on In…
FFCI: A Framework for Interpretable Automatic Evaluation of Summarization
Fajri Koto, Timothy Baldwin, Jey Han Lau
In this paper, we propose FFCI, a framework for fine-grained summarization evaluation that comprises four elements: faithfulness (degree of factual consistency with the source), fo…
Learning Contextualised Cross-lingual Word Embeddings and Alignments for Extremely Low-Resource Languages Using Parallel Corpora
Takashi Wada, Tomoharu Iwata, Yuji Matsumoto +2
We propose a new approach for learning contextualised cross-lingual word embeddings based on a small parallel corpus (e.g. a few hundred sentence pairs). Our method obtains word em…
COVID-SEE: Scientific Evidence Explorer for COVID-19 Related Research
Karin Verspoor, Simon Šuster, Yulia Otmakhova +7
We present COVID-SEE, a system for medical literature discovery based on the concept of information exploration, which builds on several distinct text analysis and natural language…
Less is More: Rejecting Unreliable Reviews for Product Question Answering
Shiwei Zhang, Xiuzhen Zhang, Jey Han Lau +2
Promptly and accurately answering questions on products is important for e-commerce applications. Manually answering product questions (e.g. on community question answering platfor…