activity
20172024
most citedIndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP

32 citations · 119 across the 25 of their papers we have counts for

collaborators

34 papers

cs.CV2024

KALE: An Artwork Image Captioning System Augmented with Heterogeneous Graph

Yanbei Jiang, Krista A. Ehinger, Jey Han Lau

Exploring the narratives conveyed by fine-art paintings is a challenge in image captioning, where the goal is to generate descriptions that not only precisely represent the visual…

cs.CL20224 cited

Not another Negation Benchmark: The NaN-NLI Test Suite for Sub-clausal Negation

Thinh Hung Truong, Yulia Otmakhova, Timothy Baldwin +3

Negation is poorly captured by current language models, although the extent of this problem is not widely understood. We introduce a natural language inference (NLI) test suite to…

cs.MM2022

Improving Visual-Semantic Embedding with Adaptive Pooling and Optimization Objective

Zijian Zhang, Chang Shu, Ya Xiao +7

Visual-Semantic Embedding (VSE) aims to learn an embedding space where related visual and semantic instances are close to each other. Recent VSE models tend to design complex struc…

cs.CL20222 cited

LED down the rabbit hole: exploring the potential of global attention for biomedical multi-document summarisation

Yulia Otmakhova, Hung Thinh Truong, Timothy Baldwin +3

In this paper we report on our submission to the Multidocument Summarisation for Literature Review (MSLR) shared task. Specifically, we adapt PRIMERA (Xiao et al., 2022) to the bio…

cs.CL20224 cited

Unsupervised Lexical Substitution with Decontextualised Embeddings

Takashi Wada, Timothy Baldwin, Yuji Matsumoto +1

We propose a new unsupervised method for lexical substitution using pre-trained language models. Compared to previous approaches that use the generative capability of language mode…

cs.CL20221 cited

Robust Task-Oriented Dialogue Generation with Contrastive Pre-training and Adversarial Filtering

Shiquan Yang, Xinting Huang, Jey Han Lau +1

Data artifacts incentivize machine learning models to learn non-transferable generalizations by taking advantage of shortcuts in the data, and there is growing evidence that data a…