Publications (85)
Characterizing AI Manipulation Risks in Brazilian YouTube Climate Discourse
Wenchao Dong, Marcelo S. Locatelli, Virgilio Almeida +1
Climate change poses a global threat to public health, food security, and economic stability. Addressing it requires evidence-based policies and a nuanced understanding of how the…
PANAS-t: A Pychometric Scale for Measuring Sentiments on Twitter
Pollyanna Gonçalves, FabrÃcio Benevenuto, Meeyoung Cha
Online social networks have become a major communication platform, where people share their thoughts and opinions about any topic real-time. The short text updates people post in t…
Prediction of Football Player Value using Bayesian Ensemble Approach
Hansoo Lee, Bayu Adhi Tama, Meeyoung Cha
The transfer fees of sports players have become astronomical. This is because bringing players of great future value to the club is essential for their survival. We present a case…
Learning Economic Indicators by Aggregating Multi-Level Geospatial Information
Sungwon Park, Sungwon Han, Donghyun Ahn +8
High-resolution daytime satellite imagery has become a promising source to study economic activities. These images display detailed terrain over large areas and allow zooming into…
Detecting Incongruity Between News Headline and Body Text via a Deep Hierarchical Encoder
Seunghyun Yoon, Kunwoo Park, Joongbo Shin +4
Some news headlines mislead readers with overrated or false information, and identifying them in advance will better assist readers in choosing proper news stories to consume. This…
Elsa: Energy-based learning for semi-supervised anomaly detection
Sungwon Han, Hyeonho Song, Seungeon Lee +2
Anomaly detection aims at identifying deviant instances from the normal data distribution. Many advances have been made in the field, including the innovative use of unsupervised c…
How You Ask Matters! Adaptive RAG Robustness to Query Variations
Yunah Jang, Megha Sundriyal, Kyomin Jung +1
Adaptive Retrieval-Augmented Generation (RAG) promises accuracy and efficiency by dynamically triggering retrieval only when needed and is widely used in practice. However, real-wo…
Fine-Grained Socioeconomic Prediction from Satellite Images with Distributional Adjustment
Donghyun Ahn, Minhyuk Song, Seungeon Lee +5
While measuring socioeconomic indicators is critical for local governments to make informed policy decisions, such measurements are often unavailable at fine-grained levels like mu…
Towards Attack-tolerant Federated Learning via Critical Parameter Analysis
Sungwon Han, Sungwon Park, Fangzhao Wu +4
Federated learning is used to train a shared model in a decentralized way without clients sharing private data with each other. Federated learning systems are susceptible to poison…
Learning Multidimensional Urban Poverty Representation with Satellite Imagery
Sungwon Park, Sumin Lee, Jihee Kim +4
Recent advances in deep learning have enabled the inference of urban socioeconomic characteristics from satellite imagery. However, models relying solely on urbanization traits oft…
EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models
Donggyu Lee, Hyeok Yun, Meeyoung Cha +3
Socio-economic causal effects depend heavily on their institutional and environmental contexts. The same intervention can produce different, even opposite, effects across regulator…
Self Supervised Vision for Climate Downscaling
Karandeep Singh, Chaeyoon Jeong, Naufal Shidqi +4
Climate change is one of the most critical challenges that our planet is facing today. Rising global temperatures are already bringing noticeable changes to Earth's weather and cli…
AI Engram: In Search of Memory Traces in Artificial Intelligence
Jea Kwon, Dong-Kyum Kim, Jiwon Kim +3
Memory formation is fundamental to intelligence, yet whether deep neural networks preserve identifiable memory traces analogous to biological memory units remains an open question.…
Self-explaining deep models with logic rule reasoning
Seungeon Lee, Xiting Wang, Sungwon Han +3
We present SELOR, a framework for integrating self-explaining capabilities into a given deep model to achieve both high prediction performance and human precision. By "human precis…
Positivity Bias in Customer Satisfaction Ratings
Kunwoo Park, Meeyoung Cha, Eunhee Rhim
Customer ratings are valuable sources to understand their satisfaction and are critical for designing better customer experiences and recommendations. The majority of customers, ho…
IO Factory: Simulating AI-Enabled Influence Campaigns at Scale
Lukasz Olejnik, Wenchao Dong, Jonas R. Kunst +4
We introduce IO Factory, an AI-driven framework for simulating information and influence campaigns as fully integrated, traceable processes. The threat of digital manipulation now…
Urban green space and happiness in developed countries
Oh-Hyun Kwon, Inho Hong, Jeasurk Yang +3
Urban green space has been regarded as contributing to citizen happiness by promoting physical and mental health. However, how urban green space and happiness are related across ma…
Cross-National Information Attacks: A Two-Decade Analysis of Troll Behavior in Korea
Jaehong Kim, Hyeonseung Kim, Jiseon Kim +4
Coordinated foreign influence operations pose a growing threat to online platforms, but detecting state-linked troll activity and tracking its evolution remain challenging. This pa…
Factuality Challenges in the Era of Large Language Models
Isabelle Augenstein, Timothy Baldwin, Meeyoung Cha +15
The emergence of tools based on Large Language Models (LLMs), such as OpenAI's ChatGPT, Microsoft's Bing Chat, and Google's Bard, has garnered immense public attention. These incre…
Classifying and Tracking International Aid Contribution Towards SDGs
Sungwon Park, Dongjoon Lee, Kyeongjin Ahn +4
International aid is a critical mechanism for promoting economic growth and well-being in developing nations, supporting progress toward the Sustainable Development Goals (SDGs). H…
Active Learning for Human-in-the-Loop Customs Inspection
Sundong Kim, Tung-Duong Mai, Sungwon Han +5
We study the human-in-the-loop customs inspection scenario, where an AI-assisted algorithm supports customs officers by recommending a set of imported goods to be inspected. If the…
Generalizable Disaster Damage Assessment via Change Detection with Vision Foundation Model
Kyeongjin Ahn, Sungwon Han, Sungwon Park +3
The increasing frequency and intensity of natural disasters call for rapid and accurate damage assessment. In response, disaster benchmark datasets from high-resolution satellite i…
Descriptive AI Ethics: Collecting and Understanding the Public Opinion
Gabriel Lima, Meeyoung Cha
There is a growing need for data-driven research efforts on how the public perceives the ethical, moral, and legal issues of autonomous AI systems. The current debate on the respon…
Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing
Jea Kwon, Jiwon Kim, Dong-kyum Kim +1
While language models remain frozen at their training state, the world evolves continuously. Knowledge editing has emerged as a key alternative to full retraining, but its deployme…
Semi-supervised Shelter Mapping for WASH Accessibility Assessment in Rohingya Refugee Camps
Kyeongjin Ahn, YongHun Suh, Sungwon Han +3
Lack of access to Water, Sanitation, and Hygiene (WASH) services is a major public health concern in refugee camps, where extreme crowding accelerates the spread of communicable di…
FedMID: A Data-Free Method for Using Intermediate Outputs as a Defense Mechanism Against Poisoning Attacks in Federated Learning
Sungwon Han, Hyeonho Song, Sungwon Park +1
Federated learning combines local updates from clients to produce a global model, which is susceptible to poisoning attacks. Most previous defense strategies relied on vectors deri…
Adversarial Style Augmentation via Large Language Model for Robust Fake News Detection
Sungwon Park, Sungwon Han, Xing Xie +2
The spread of fake news harms individuals and presents a critical social challenge that must be addressed. Although numerous algorithmic and insightful features have been developed…
Knowledge Sharing via Domain Adaptation in Customs Fraud Detection
Sungwon Park, Sundong Kim, Meeyoung Cha
Knowledge of the changing traffic is critical in risk management. Customs offices worldwide have traditionally relied on local resources to accumulate knowledge and detect tax frau…
Mining the Minds of Customers from Online Chat Logs
Kunwoo Park, Jaewoo Kim, Jaram Park +4
This study investigates factors that may determine satisfaction in customer service operations. We utilized more than 170,000 online chat sessions between customers and agents to i…
Lightweight and Robust Representation of Economic Scales from Satellite Imagery
Sungwon Han, Donghyun Ahn, Hyunji Cha +3
Satellite imagery has long been an attractive data source that provides a wealth of information on human-inhabited areas. While super resolution satellite images are rapidly becomi…
Blaming Humans and Machines: What Shapes People's Reactions to Algorithmic Harm
Gabriel Lima, Nina GrgiÄ-HlaÄa, Meeyoung Cha
Artificial intelligence (AI) systems can cause harm to people. This research examines how individuals react to such harm through the lens of blame. Building upon research suggestin…
Classification of Goods Using Text Descriptions With Sentences Retrieval
Eunji Lee, Sundong Kim, Sihyun Kim +8
The task of assigning and validating internationally accepted commodity code (HS code) to traded goods is one of the critical functions at the customs office. This decision is cruc…
SQuARe: A Large-Scale Dataset of Sensitive Questions and Acceptable Responses Created Through Human-Machine Collaboration
Hwaran Lee, Seokhee Hong, Joonsuk Park +10
The potential social harms that large language models pose, such as generating offensive content and reinforcing biases, are steeply rising. Existing works focus on coping with thi…
Dawn of the Selfie Era: The Whos, Wheres, and Hows of Selfies on Instagram
Flávio Souza, Diego de Las Casas, VinÃcius Flores +4
Online interactions are increasingly involving images, especially those containing human faces, which are naturally attention grabbing and more effective at conveying feelings than…
Misinformation, Believability, and Vaccine Acceptance Over 40 Countries: Takeaways From the Initial Phase of The COVID-19 Infodemic
Karandeep Singh, Gabriel Lima, Meeyoung Cha +4
The COVID-19 pandemic has been damaging to the lives of people all around the world. Accompanied by the pandemic is an infodemic, an abundant and uncontrolled spreading of potentia…
A Comprehensive Approach to Unsupervised Embedding Learning based on AND Algorithm
Sungwon Han, Yizhan Xu, Sungwon Park +2
Unsupervised embedding learning aims to extract good representation from data without the need for any manual labels, which has been a critical challenge in many supervised learnin…
Textual Supervision Enhances Geospatial Representations in Vision-Language Models
Marcelo Sartori Locatelli, Fernando Tonucci, Jea Kwon +5
Geospatial understanding is a critical yet underexplored dimension in the development of machine learning systems for tasks such as image geolocation and spatial reasoning. In this…
Poverty mapping in Mongolia with AI-based Ger detection reveals urban slums persist after the COVID-19 pandemic
Jeasurk Yang, Sumin Lee, Sungwon Park +2
Mongolia is among the countries undergoing rapid urbanization, and its temporary nomadic dwellings-known as Ger-have expanded into urban areas. Ger settlements in cities are increa…
GraphFC: Customs Fraud Detection with Label Scarcity
Karandeep Singh, Yu-Che Tsai, Cheng-Te Li +2
Custom officials across the world encounter huge volumes of transactions. With increased connectivity and globalization, the customs transactions continue to grow every year. Assoc…
DualFair: Fair Representation Learning at Both Group and Individual Levels via Contrastive Self-supervision
Sungwon Han, Seungeon Lee, Fangzhao Wu +5
Algorithmic fairness has become an important machine learning problem, especially for mission-critical Web applications. This work presents a self-supervised model, called DualFair…
Modeling the adoption of innovations in the presence of geographic and media influences
Jameson L. Toole, Meeyoung Cha, Marta C. Gonzalez
While there has been much work examining the affects of social network structure on innovation adoption, models to date have lacked important features such as meta-populations refl…
Fashion Conversation Data on Instagram
Yu-I Ha, Sejeong Kwon, Meeyoung Cha +1
The fashion industry is establishing its presence on a number of visual-centric social media like Instagram. This creates an interesting clash as fashion brands that have tradition…
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
Jea Kwon, Luiz Felipe Vecchietti, Sungwon Park +1
Humans display significant uncertainty when confronted with moral dilemmas, yet the extent of such uncertainty in machines and AI agents remains underexplored. Recent studies have…
Machine Behavior in Relational Moral Dilemmas: Moral Rightness, Predicted Human Behavior, and Model Decisions
Jiseon Kim, Jea Kwon, Luiz Felipe Vecchietti +3
Human moral judgment is context-dependent and modulated by interpersonal relationships. As large language models (LLMs) increasingly function as decision-support systems, determini…
I Am Not Them: Fluid Identities and Persistent Out-group Bias in Large Language Models
Wenchao Dong, Assem Zhunis, Hyojin Chin +2
We explored cultural biases-individualism vs. collectivism-in ChatGPT across three Western languages (i.e., English, German, and French) and three Eastern languages (i.e., Chinese,…
GeoX: Mastering Geospatial Reasoning Through Self-Play and Verifiable Rewards
Kyeongjin Ahn, Seungeon Lee, Krishna P. Gummadi +1
Geospatial reasoning requires solving image-grounded problems over the complex spatial structure of a scene. However, developing this capability is hindered by the cost of annotati…
Explainable Product Classification for Customs
Eunji Lee, Sihyeon Kim, Sundong Kim +3
The task of assigning internationally accepted commodity codes (aka HS codes) to traded goods is a critical function of customs offices. Like court decisions made by judges, this t…
Bilinear representation mitigates reversal curse and enables consistent model editing
Dong-Kyum Kim, Minsung Kim, Jea Kwon +2
The reversal curse--a language model's inability to infer an unseen fact "B is A" from a learned fact "A is B"--is widely considered a fundamental limitation. We show that this is…
How malicious AI swarms can threaten democracy: The fusion of agentic AI and LLMs marks a new frontier in information warfare
Daniel Thilo Schroeder, Meeyoung Cha, Andrea Baronchelli +19
Advances in AI offer the prospect of manipulating beliefs and behaviors on a population-wide level. Large language models and autonomous agents now let influence campaigns reach un…
Human Perceptions on Moral Responsibility of AI: A Case Study in AI-Assisted Bail Decision-Making
Gabriel Lima, Nina GrgiÄ-HlaÄa, Meeyoung Cha
How to attribute responsibility for autonomous artificial intelligence (AI) systems' actions has been widely debated across the humanities and social science disciplines. This work…
Mapping Emerging Climate Misinformation Playbooks in the Global South
Marcelo Sartori Locatelli, Wenchao Dong, Pedro Loures Alzamora +4
Climate misinformation continues to erode support for climate action, a challenge that is especially acute in the Global South, where high climate vulnerability intersects with dev…
ERA: Evidence-based Reliability Alignment for Honest Retrieval-Augmented Generation
Sunguk Shin, Meeyoung Cha, Byung-Jun Lee +1
Retrieval-Augmented Generation (RAG) grounds language models in factual evidence but introduces critical challenges regarding knowledge conflicts between internalized parameters an…
A Comparative Study of Reference Reliability in Multiple Language Editions of Wikipedia
Aitolkyn Baigutanova, Diego Saez-Trumper, Miriam Redi +2
Information presented in Wikipedia articles must be attributable to reliable published sources in the form of references. This study examines over 5 million Wikipedia articles to a…
GeoReg: Weight-Constrained Few-Shot Regression for Socio-Economic Estimation using LLM
Kyeongjin Ahn, Sungwon Han, Seungeon Lee +6
Socio-economic indicators like regional GDP, population, and education levels, are crucial to shaping policy decisions and fostering sustainable development. This research introduc…
Detecting Contextomized Quotes in News Headlines by Contrastive Learning
Seonyeong Song, Hyeonho Song, Kunwoo Park +2
Quotes are critical for establishing credibility in news articles. A direct quote enclosed in quotation marks has a strong visual appeal and is a sign of a reliable citation. Unfor…
Erase or Hide? Suppressing Spurious Unlearning Neurons for Robust Unlearning
Nakyeong Yang, Dong-Kyum Kim, Jea Kwon +3
Large language models trained on web-scale data can memorize private or sensitive knowledge, raising significant privacy risks. Although some unlearning methods mitigate these risk…
Quantitative Analysis of Cultural Dynamics Seen from an Event-based Social Network
Bayu Adhi Tama, Jaehong Kim, Jaehyuk Park +2
Culture is a collection of connected and potentially interactive patterns that characterize a social group or a passed-on idea that people acquire as members of society. While offl…
"Trust me, I have a Ph.D.": A Propensity Score Analysis on the Halo Effect of Disclosing One's Offline Social Status in Online Communities
Kunwoo Park, Haewoon Kwak, Hyunho Song +1
Online communities adopt various reputation schemes to measure content quality. This study analyzes the effect of a new reputation scheme that exposes one's offline social status,…
Disruption in the Chinese E-Commerce During COVID-19
Yuan Yuan, Muzhi Guan, Zhilun Zhou +4
The recent outbreak of the novel coronavirus (COVID-19) has infected millions of citizens worldwide and claimed many lives. This paper examines its impact on the Chinese e-commerce…
FedDefender: Client-Side Attack-Tolerant Federated Learning
Sungwon Park, Sungwon Han, Fangzhao Wu +4
Federated learning enables learning from decentralized data sources without compromising privacy, which makes it a crucial technique. However, it is vulnerable to model poisoning a…
BaitWatcher: A lightweight web interface for the detection of incongruent news headlines
Kunwoo Park, Taegyun Kim, Seunghyun Yoon +2
In digital environments where substantial amounts of information are shared online, news headlines play essential roles in the selection and diffusion of news articles. Some news a…
Generalizable Slum Detection from Satellite Imagery with Mixture-of-Experts
Sumin Lee, Sungwon Park, Jeasurk Yang +2
Satellite-based slum segmentation holds significant promise in generating global estimates of urban poverty. However, the morphological heterogeneity of informal settlements presen…
How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
Minsung Kim, Dong-Kyum Kim, Jea Kwon +3
Large language models leverage both parametric knowledge acquired during pretraining and in-context knowledge provided at inference time. Crucially, when these sources conflict, mo…
Dual-Margin Embedding for Fine-Grained Long-Tailed Plant Taxonomy
Cheng Yaw Low, Heejoon Koo, Jaewoo Park +1
Taxonomic classification of ecological families, genera, and species underpins biodiversity monitoring and conservation. Existing computer vision methods typically address fine-gra…
Robust Optimization in Protein Fitness Landscapes Using Reinforcement Learning in Latent Space
Minji Lee, Luiz Felipe Vecchietti, Hyunkyu Jung +3
Proteins are complex molecules responsible for different functions in nature. Enhancing the functionality of proteins and cellular fitness can significantly impact various industri…
Achievement and Friends: Key Factors of Player Retention Vary Across Player Levels in Online Multiplayer Games
Kunwoo Park, Meeyoung Cha, Haewoon Kwak +1
Retaining players over an extended period of time is a long-standing challenge in game industry. Significant effort has been paid to understanding what motivates players enjoy game…
GeoSEE: Regional Socio-Economic Estimation With a Large Language Model
Sungwon Han, Donghyun Ahn, Seungeon Lee +5
Moving beyond traditional surveys, combining heterogeneous data sources with AI-driven inference models brings new opportunities to measure socio-economic conditions, such as pover…
Longitudinal Assessment of Reference Quality on Wikipedia
Aitolkyn Baigutanova, Jaehyeon Myung, Diego Saez-Trumper +4
Wikipedia plays a crucial role in the integrity of the Web. This work analyzes the reliability of this global encyclopedia through the lens of its references. We operationalize the…
Persistent Sharing of Fitness App Status on Twitter
Kunwoo Park, Ingmar Weber, Meeyoung Cha +1
As the world becomes more digitized and interconnected, information that was once considered to be private such as one's health status is now being shared publicly. To understand t…
Comparing and Combining Sentiment Analysis Methods
Pollyanna Gonçalves, Matheus Araújo, FabrÃcio Benevenuto +1
Several messages express opinions about events, products, and services, political views or even their author's emotional state and mood. Sentiment analysis has been used in several…
Recent advances in interpretable machine learning using structure-based protein representations
Luiz Felipe Vecchietti, Minji Lee, Begench Hangeldiyev +5
Recent advancements in machine learning (ML) are transforming the field of structural biology. For example, AlphaFold, a groundbreaking neural network for protein structure predict…
Persona Setting Pitfall: Persistent Outgroup Biases in Large Language Models Arising from Social Identity Adoption
Wenchao Dong, Assem Zhunis, Dongyoung Jeong +3
Drawing parallels between human cognition and artificial intelligence, we explored how large language models (LLMs) internalize identities imposed by targeted prompts. Informed by…
Responsible AI and Its Stakeholders
Gabriel Lima, Meeyoung Cha
Responsible Artificial Intelligence (AI) proposes a framework that holds all stakeholders involved in the development of AI to be responsible for their systems. It, however, fails…
Collecting the Public Perception of AI and Robot Rights
Gabriel Lima, Changyeon Kim, Seungho Ryu +2
Whether to give rights to artificial intelligence (AI) and robots has been a sensitive topic since the European Parliament proposed advanced robots could be granted "electronic per…
Improving Unsupervised Image Clustering With Robust Learning
Sungwon Park, Sungwon Han, Sundong Kim +4
Unsupervised image clustering methods often introduce alternative objectives to indirectly train the model and are subject to faulty predictions and overconfident results. To overc…
Retrieval Augmented Time Series Forecasting
Sungwon Han, Seungeon Lee, Meeyoung Cha +2
Time series forecasting uses historical data to predict future trends, leveraging the relationships between past observations and available features. In this paper, we propose RAFT…
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
Jiseon Kim, Jea Kwon, Luiz Felipe Vecchietti +2
Deploying large language models (LLMs) with agency in real-world applications raises critical questions about how these models will behave. In particular, how will their decisions…
COVID-19 Vaccine Acceptance in the US and UK in the Early Phase of the Pandemic: AI-Generated Vaccines Hesitancy for Minors, and the Role of Governments
Gabriel Lima, Meeyoung Cha, Chiyoung Cha +1
This study presents survey results of the public's willingness to get vaccinated against COVID-19 during an early phase of the pandemic and examines factors that could influence va…
Social Bootstrapping: How Pinterest and Last.fm Social Communities Benefit by Borrowing Links from Facebook
Changtao Zhong, Mostafa Salehi, Sunil Shah +3
How does one develop a new online community that is highly engaging to each user and promotes social interaction? A number of websites offer friend-finding features that help users…
FedX: Unsupervised Federated Learning with Cross Knowledge Distillation
Sungwon Han, Sungwon Park, Fangzhao Wu +4
This paper presents FedX, an unsupervised federated learning framework. Our model learns unbiased representation from decentralized and heterogeneous local data. It employs a two-s…
Puppets or partners? Governing cyborg propaganda in the digital public square
Jonas R. Kunst, Kinga Bierwiaczonek, Meeyoung Cha +12
The distinction between genuine grassroots activism and automated influence operations is collapsing. While contemporary policy debates prioritize fully autonomous generative agent…
The Conflict Between Explainable and Accountable Decision-Making Algorithms
Gabriel Lima, Nina GrgiÄ-HlaÄa, Jin Keun Jeong +1
Decision-making algorithms are being used in important decisions, such as who should be enrolled in health care programs and be hired. Even though these systems are currently deplo…
Risk Communication in Asian Countries: COVID-19 Discourse on Twitter
Sungkyu Park, Sungwon Han, Jeongwook Kim +6
COVID-19 has become one of the most widely talked about topics on social media. This research characterizes risk communication patterns by analyzing the public discourse on the nov…
Evaluating the Robustness of Trigger Set-Based Watermarks Embedded in Deep Neural Networks
Suyoung Lee, Wonho Song, Suman Jana +2
Trigger set-based watermarking schemes have gained emerging attention as they provide a means to prove ownership for deep neural network model owners. In this paper, we argue that…
The Conflict Between People's Urge to Punish AI and Legal Systems
Gabriel Lima, Meeyoung Cha, Chihyung Jeon +1
Regulating artificial intelligence (AI) has become necessary in light of its deployment in high-risk scenarios. This paper explores the proposal to extend legal personhood to AI an…