papers

Publications (11)

cs.CL2021

Evaluation of BERT and ALBERT Sentence Embedding Performance on Downstream NLP Tasks

Hyunjin Choi, Judong Kim, Seongho Joe +1

Contextualized representations from a pre-trained language model are central to achieve a high performance on downstream NLP task. The pre-trained BERT and A Lite BERT (ALBERT) mod…

cs.CL2023

Shuffle & Divide: Contrastive Learning for Long Text

Joonseok Lee, Seongho Joe, Kyoungwon Park +4

We propose a self-supervised learning method for long text documents based on contrastive learning. A key to our method is Shuffle and Divide (SaD), a simple text augmentation algo…

cs.LG2023

Is Cross-modal Information Retrieval Possible without Training?

Hyunjin Choi, Hyunjae Lee, Seongho Joe +1

Encoded representations from a pretrained deep learning model (e.g., BERT text embeddings, penultimate CNN layer activations of an image) convey a rich set of features beneficial f…

cs.CL2021

Analyzing Zero-shot Cross-lingual Transfer in Supervised NLP Tasks

Hyunjin Choi, Judong Kim, Seongho Joe +2

In zero-shot cross-lingual transfer, a supervised NLP task trained on a corpus in one language is directly applicable to another language without any additional training. A source…

cs.CL2025

Correcting Negative Bias in Large Language Models through Negative Attention Score Alignment

Sangwon Yu, Jongyoon Song, Bongkyu Hwang +7

A binary decision task, like yes-no questions or answer verification, reflects a significant real-world scenario such as where users look for confirmation about the correctness of…

cs.CL2021

KoreALBERT: Pretraining a Lite BERT Model for Korean Language Understanding

Hyunjae Lee, Jaewoong Yoon, Bonggyu Hwang +3

A Lite BERT (ALBERT) has been introduced to scale up deep bidirectional representation learning for natural languages. Due to the lack of pretrained ALBERT models for Korean langua…

cs.CL2024

Entity-level Factual Adaptiveness of Fine-tuning based Abstractive Summarization Models

Jongyoon Song, Nohil Park, Bongkyu Hwang +4

Abstractive summarization models often generate factually inconsistent content particularly when the parametric knowledge of the model conflicts with the knowledge in the input doc…

cs.CV2023

ContraCluster: Learning to Classify without Labels by Contrastive Self-Supervision and Prototype-Based Semi-Supervision

Seongho Joe, Byoungjip Kim, Hoyoung Kang +5

The recent advances in representation learning inspire us to take on the challenging problem of unsupervised image classification tasks in a principled way. We propose ContraCluste…

cs.CL2022

Enhancing Semantic Understanding with Self-supervised Methods for Abstractive Dialogue Summarization

Hyunjae Lee, Jaewoong Yun, Hyunjin Choi +2

Contextualized word embeddings can lead to state-of-the-art performances in natural language understanding. Recently, a pre-trained deep contextualized text encoder such as BERT ha…

cs.LG2021

SelfMatch: Combining Contrastive Self-Supervision and Consistency for Semi-Supervised Learning

Byoungjip Kim, Jinho Choo, Yeong-Dae Kwon +3

This paper introduces SelfMatch, a semi-supervised learning method that combines the power of contrastive self-supervised learning and consistency regularization. SelfMatch consist…

cs.CV2021

BiHPF: Bilateral High-Pass Filters for Robust Deepfake Detection

Yonghyun Jeong, Doyeon Kim, Seungjai Min +3

The advancement in numerous generative models has a two-fold effect: a simple and easy generation of realistic synthesized images, but also an increased risk of malicious abuse of…