papers

Publications (69)

cs.SI2024

A Survey on the Role of Crowds in Combating Online Misinformation: Annotators, Evaluators, and Creators

Bing He, Yibo Hu, Yeon-Chang Lee +3

Online misinformation poses a global risk with significant real-world consequences. To combat misinformation, current research relies on professionals like journalists and fact-che…

cs.CL2024

MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms

Yiqiao Jin, Minje Choi, Gaurav Verma +2

Social media platforms are hubs for multimodal information exchange, encompassing text, images, and videos, making it challenging for machines to comprehend the information or emot…

hep-th2024

One point functions in large vector models at finite chemical potential

Justin R. David, Srijan Kumar

We evaluate the thermal one point function of higher spin currents in the critical model of complex scalars interacting with a quartic potential and the Gross-Neveu m…

cs.SI2017

An Army of Me: Sockpuppets in Online Discussion Communities

Srijan Kumar, Justin Cheng, Jure Leskovec +1

In online discussion communities, users can interact and share information and opinions on a wide variety of topics. However, some users may create multiple identities, or sockpupp…

cs.LG2024

Empowering Interdisciplinary Insights with Dynamic Graph Embedding Trajectories

Yiqiao Jin, Andrew Zhao, Yeon-Chang Lee +3

We developed DyGETViz, a novel framework for effectively visualizing dynamic graphs (DGs) that are ubiquitous across diverse real-world systems. This framework leverages recent adv…

cs.MA2025

Topological Structure Learning Should Be A Research Priority for LLM-Based Multi-Agent Systems

Jiaxi Yang, Mengqi Zhang, Yiqiao Jin +8

Large Language Model-based Multi-Agent Systems (MASs) have emerged as a powerful paradigm for tackling complex tasks through collaborative intelligence. However, the topology of th…

cs.IR2024

Adversarial Text Rewriting for Text-aware Recommender Systems

Sejoon Oh, Gaurav Verma, Srijan Kumar

Text-aware recommender systems incorporate rich textual features, such as titles and descriptions, to generate item recommendations for users. The use of textual features helps mit…

cs.SI2022

Characterizing, Detecting, and Predicting Online Ban Evasion

Manoj Niverthi, Gaurav Verma, Srijan Kumar

Moderators and automated methods enforce bans on malicious users who engage in disruptive behavior. However, malicious users can easily create a new account to evade such bans. Pre…

cs.SI2023

Representation Learning in Continuous-Time Dynamic Signed Networks

Kartik Sharma, Mohit Raghavendra, Yeon Chang Lee +2

Signed networks allow us to model conflicting relationships and interactions, such as friend/enemy and support/oppose. These signed interactions happen in real-time. Modeling such…

cs.LG2022

Robustness of Fusion-based Multimodal Classifiers to Cross-Modal Content Dilutions

Gaurav Verma, Vishwa Vinay, Ryan A. Rossi +1

As multimodal learning finds applications in a wide variety of high-stakes societal tasks, investigating their robustness becomes important. Existing work has focused on understand…

cs.CL2025

MedHalu: Hallucinations in Responses to Healthcare Queries by Large Language Models

Vibhor Agarwal, Yiqiao Jin, Mohit Chandra +3

Large language models (LLMs) are starting to complement traditional information seeking mechanisms such as web search. LLM-powered chatbots like ChatGPT are gaining prominence amon…

cs.CL2024

A Community-Centric Perspective for Characterizing and Detecting Anti-Asian Violence-Provoking Speech

Gaurav Verma, Rynaa Grover, Jiawei Zhou +4

Violence-provoking speech -- speech that implicitly or explicitly promotes violence against the members of the targeted community, contributed to a massive surge in anti-Asian crim…

hep-th2025

High to low temperature: model at large

Justin R. David, Srijan Kumar

We study the vector model for scalars with quartic interaction at large on without the singlet constraint. The non-trivial fixed point of the model is de…

cs.CL2026

Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings

Kartik Sharma, Yiqiao Jin, Rakshit Trivedi +1

Large language models (LLMs) acquire knowledge across diverse domains such as science, history, and geography encountered during generative pre-training. However, due to their stoc…

cs.CL2024

Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space

Gaurav Verma, Minje Choi, Kartik Sharma +3

Multimodal large language models (MLLMs) like LLaVA and GPT-4(V) enable general-purpose conversations about images with the language modality. As off-the-shelf MLLMs may have limit…

cs.SI2018

False Information on Web and Social Media: A Survey

Srijan Kumar, Neil Shah

False information can be created and spread easily through the web and social media platforms, resulting in widespread real-world impact. Characterizing how false information proli…

cs.CL2025

UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models

Sejoon Oh, Yiqiao Jin, Megha Sharma +4

Multimodal large language models (MLLMs) have revolutionized vision-language understanding but remain vulnerable to multimodal jailbreak attacks, where adversarial inputs are metic…

cs.SI2023

Characterizing and Predicting Social Correction on Twitter

Yingchen Ma, Bing He, Nathan Subrahmanian +1

Online misinformation has been a serious threat to public health and society. Social media users are known to reply to misinformation posts with counter-misinformation messages, wh…

cs.SI2015

VEWS: A Wikipedia Vandal Early Warning System

Srijan Kumar, Francesca Spezzano, V. S. Subrahmanian

We study the problem of detecting vandals on Wikipedia before any human or known vandalism detection system reports flagging potential vandals so that such users can be presented e…

cs.SI2021

Racism is a Virus: Anti-Asian Hate and Counterspeech in Social Media during the COVID-19 Crisis

Bing He, Caleb Ziems, Sandeep Soni +3

The spread of COVID-19 has sparked racism and hate on social media targeted towards Asian communities. However, little is known about how racial hate spreads during a pandemic and…

cs.CL2026

UniSD: Towards a Unified Self-Distillation Framework for Large Language Models

Yiqiao Jin, Yiyang Wang, Lucheng Fu +7

Self-distillation (SD) offers a promising path for adapting large language models (LLMs) without relying on stronger external teachers. However, SD in autoregressive LLMs remains c…

cs.SI2025

The Influence of Text Variation on User Engagement in Cross-Platform Content Sharing

Yibo Hu, Yiqiao Jin, Meng Ye +2

In today's cross-platform social media landscape, understanding factors that drive engagement for multimodal content, especially text paired with visuals, remains complex. This stu…

hep-th2023

Thermal one point functions, large and interior geometry of black holes

Justin R. David, Srijan Kumar

We study thermal one point functions of massive scalars in black holes. These are induced by coupling the scalar to either the Weyl tensor squared or the Gauss-Bonnet t…

cs.LG2021

PETGEN: Personalized Text Generation Attack on Deep Sequence Embedding-based Classification Models

Bing He, Mustaque Ahamad, Srijan Kumar

What should a malicious user write next to fool a detection model? Identifying malicious users is critical to ensure the safety and integrity of internet platforms. Several deep le…

hep-th2025

The large vector model on

Justin R. David, Srijan Kumar

We develop a method to evaluate the partition function and energy density of a massive scalar on a 2-sphere of radius and at finite temperature as power series in $\fracβ…

cs.SI2024

Towards Fair Graph Anomaly Detection: Problem, Benchmark Datasets, and Evaluation

Neng Kai Nigel Neo, Yeon-Chang Lee, Yiqiao Jin +2

The Fair Graph Anomaly Detection (FairGAD) problem aims to accurately detect anomalous nodes in an input graph while avoiding biased predictions against individuals from sensitive…

cs.LG2021

Influence-guided Data Augmentation for Neural Tensor Completion

Sejoon Oh, Sungchul Kim, Ryan A. Rossi +1

How can we predict missing values in multi-dimensional data (or tensors) more accurately? The task of tensor completion is crucial in many applications such as personalized recomme…

cs.SI2020

Higher-Order Label Homogeneity and Spreading in Graphs

Dhivya Eswaran, Srijan Kumar, Christos Faloutsos

Do higher-order network structures aid graph semi-supervised learning? Given a graph and a few labeled vertices, labeling the remaining vertices is a high-impact problem with appli…

cs.CL2024

Introducing v0.5 of the AI Safety Benchmark from MLCommons

Bertie Vidgen, Adarsh Agrawal, Ahmed M. Ahmed +97

This paper introduces v0.5 of the AI Safety Benchmark, which has been created by the MLCommons AI Safety Working Group. The AI Safety Benchmark has been designed to assess the safe…

cs.CL2023

Cross-Modal Attribute Insertions for Assessing the Robustness of Vision-and-Language Learning

Shivaen Ramshetty, Gaurav Verma, Srijan Kumar

The robustness of multimodal deep learning models to realistic changes in the input text is critical for their applicability to important tasks such as text-to-image retrieval and…

cs.SI2020

The Role of the Crowd in Countering Misinformation: A Case Study of the COVID-19 Infodemic

Nicholas Micallef, Bing He, Srijan Kumar +2

Fact checking by professionals is viewed as a vital defense in the fight against misinformation.While fact checking is important and its impact has been significant, fact checks co…

cs.CL2023

Better to Ask in English: Cross-Lingual Evaluation of Large Language Models for Healthcare Queries

Yiqiao Jin, Mohit Chandra, Gaurav Verma +3

Large language models (LLMs) are transforming the ways the general public accesses and consumes information. Their influence is particularly pronounced in pivotal sectors like heal…

hep-th2023

Thermal one-point functions: CFT's with fermions, large and large spin

Justin R. David, Srijan Kumar

We apply the OPE inversion formula on thermal two-point functions of fermions to obtain thermal one-point function of fermion bi-linears appearing in the corresponding OPE. We prim…

cs.CL2015

Linguistic Harbingers of Betrayal: A Case Study on an Online Strategy Game

Vlad Niculae, Srijan Kumar, Jordan Boyd-Graber +1

Interpersonal relations are fickle, with close friendships often dissolving into enmity. In this work, we explore linguistic cues that presage such transitions by studying dyadic i…

cs.SI2024

Supporters and Skeptics: LLM-based Analysis of Engagement with Mental Health (Mis)Information Content on Video-sharing Platforms

Viet Cuong Nguyen, Mini Jain, Abhijat Chauhan +7

Over one in five adults in the US lives with a mental illness. In the face of a shortage of mental health professionals and offline resources, online short-form video content has g…

cs.SI2017

FairJudge: Trustworthy User Prediction in Rating Platforms

Srijan Kumar, Bryan Hooi, Disha Makhija +3

Rating platforms enable large-scale collection of user opinion about items (products, other users, etc.). However, many untrustworthy users give fraudulent ratings for excessive mo…

cs.CL2023

Factify 2: A Multimodal Fake News and Satire News Dataset

S Suryavardan, Shreyash Mishra, Parth Patwa +9

The internet gives the world an open platform to express their views and share their stories. While this is very valuable, it makes fake news one of our society's most pressing pro…

cs.CL2023

Overview of Memotion 3: Sentiment and Emotion Analysis of Codemixed Hinglish Memes

Shreyash Mishra, S Suryavardan, Megha Chakraborty +9

Analyzing memes on the internet has emerged as a crucial endeavor due to the impact this multi-modal form of content wields in shaping online discourse. Memes have become a powerfu…

cs.LG2025

ROBAD: Robust Adversary-aware Local-Global Attended Bad Actor Detection Sequential Model

Bing He, Mustaque Ahamad, Srijan Kumar

Detecting bad actors is critical to ensure the safety and integrity of internet platforms. Several deep learning-based models have been developed to identify such users. These mode…

cs.IR2022

M2TRec: Metadata-aware Multi-task Transformer for Large-scale and Cold-start free Session-based Recommendations

Walid Shalaby, Sejoon Oh, Amir Afsharinejad +2

Session-based recommender systems (SBRSs) have shown superior performance over conventional methods. However, they show limited scalability on large-scale industrial datasets since…

cs.LG2025

Personalized Layer Selection for Graph Neural Networks

Kartik Sharma, Vineeth Rakesh, Yingtong Dou +2

Graph Neural Networks (GNNs) combine node attributes over a fixed granularity of the local graph structure around a node to predict its label. However, different nodes may relate t…

cs.SI2018

Learning Dynamic Embeddings from Temporal Interactions

Srijan Kumar, Xikun Zhang, Jure Leskovec

Modeling a sequence of interactions between users and items (e.g., products, posts, or courses) is crucial in domains such as e-commerce, social networking, and education to predic…

cs.SI2018

Community Interaction and Conflict on the Web

Srijan Kumar, William L. Hamilton, Jure Leskovec +1

Users organize themselves into communities on web platforms. These communities can interact with one another, often leading to conflicts and toxic interactions. However, little is…

cs.CV2021

M2P2: Multimodal Persuasion Prediction using Adaptive Fusion

Chongyang Bai, Haipeng Chen, Srijan Kumar +2

Identifying persuasive speakers in an adversarial environment is a critical task. In a national election, politicians would like to have persuasive speakers campaign on their behal…

cs.CL2025

SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression

Yiqiao Jin, Kartik Sharma, Vineeth Rakesh +4

Retrieval-augmented Generation (RAG) extends large language models (LLMs) with external knowledge but faces key challenges: restricted effective context length and redundancy in re…

cs.SI2023

Predicting Information Pathways Across Online Communities

Yiqiao Jin, Yeon-Chang Lee, Kartik Sharma +4

The problem of community-level information pathway prediction (CLIPP) aims at predicting the transmission trajectory of content across online communities. A successful solution to…

cs.SI2021

Deception Detection in Group Video Conversations using Dynamic Interaction Networks

Srijan Kumar, Chongyang Bai, V. S. Subrahmanian +1

Detecting groups of people who are jointly deceptive in video conversations is crucial in settings such as meetings, sales pitches, and negotiations. Past work on deception in vide…

cs.SI2019

Predicting Dynamic Embedding Trajectory in Temporal Interaction Networks

Srijan Kumar, Xikun Zhang, Jure Leskovec

Modeling sequential interactions between users and items/products is crucial in domains such as e-commerce, social networking, and education. Representation learning presents an at…

cs.LG2022

Overcoming Language Disparity in Online Content Classification with Multimodal Learning

Gaurav Verma, Rohit Mujumdar, Zijie J. Wang +2

Advances in Natural Language Processing (NLP) have revolutionized the way researchers and practitioners address crucial societal problems. Large language models are now the standar…

cs.SI2023

Reinforcement Learning-based Counter-Misinformation Response Generation: A Case Study of COVID-19 Vaccine Misinformation

Bing He, Mustaque Ahamad, Srijan Kumar

The spread of online misinformation threatens public health, democracy, and the broader society. While professional fact-checkers form the first line of defense by fact-checking po…

cs.SI2022

Graph Vulnerability and Robustness: A Survey

Scott Freitas, Diyi Yang, Srijan Kumar +2

The study of network robustness is a critical tool in the characterization and sense making of complex interconnected systems such as infrastructure, communication and social netwo…

cs.AI2026

Sysformer: Safeguarding Frozen Large Language Models with Adaptive System Prompts

Kartik Sharma, Yiqiao Jin, Vineeth Rakesh +4

As large language models (LLMs) are deployed in safety-critical settings, it is essential to ensure that their responses comply with safety standards. Prior research has revealed t…

cs.CL2026

SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding

Yiqiao Jin, Rachneet Kaur, Zhen Zeng +2

Multi-page visual documents such as manuals, brochures, presentations, and posters convey key information through layout, colors, icons, and cross-slide references. While multimoda…

cs.CL2025

Do Large Language Models Align with Core Mental Health Counseling Competencies?

Viet Cuong Nguyen, Mohammad Taher, Dongwan Hong +8

The rapid evolution of Large Language Models (LLMs) presents a promising solution to the global shortage of mental health professionals. However, their alignment with essential cou…

cs.IR2022

Implicit Session Contexts for Next-Item Recommendations

Sejoon Oh, Ankur Bhardwaj, Jongseok Han +3

Session-based recommender systems capture the short-term interest of a user within a session. Session contexts (i.e., a user's high-level interests or intents within a session) are…

cs.CL2016

Citation Classification for Behavioral Analysis of a Scientific Field

David Jurgens, Srijan Kumar, Raine Hoover +2

Citations are an important indicator of the state of a scientific field, reflecting how authors frame their work, and influencing uptake by future scholars. However, our understand…

cs.IR2024

FINEST: Stabilizing Recommendations by Rank-Preserving Fine-Tuning

Sejoon Oh, Berk Ustun, Julian McAuley +1

Modern recommender systems may output considerably different recommendations due to small perturbations in the training data. Changes in the data from a single user will alter the…

cs.CL2023

Adversarial Robustness of Prompt-based Few-Shot Learning for Natural Language Understanding

Venkata Prabhakara Sarath Nookala, Gaurav Verma, Subhabrata Mukherjee +1

State-of-the-art few-shot learning (FSL) methods leverage prompt-based fine-tuning to obtain remarkable results for natural language understanding (NLU) tasks. While much of the pr…

cs.AI2025

A Framework for Situating Innovations, Opportunities, and Challenges in Advancing Vertical Systems with Large AI Models

Gaurav Verma, Jiawei Zhou, Mohit Chandra +2

Large artificial intelligence (AI) models have garnered significant attention for their remarkable, often "superhuman", performance on standardized benchmarks. However, when these…

cs.SI2024

Corrective or Backfire: Characterizing and Predicting User Response to Social Correction

Bing He, Yingchen Ma, Mustaque Ahamad +1

Online misinformation poses a global risk with harmful implications for society. Ordinary social media users are known to actively reply to misinformation posts with counter-misinf…

cs.IR2022

Rank List Sensitivity of Recommender Systems to Interaction Perturbations

Sejoon Oh, Berk Ustun, Julian McAuley +1

Prediction models can exhibit sensitivity with respect to training data: small changes in the training data can produce models that assign conflicting predictions to individual dat…

cs.AI2026

AgentArk: Distilling Multi-Agent Intelligence into a Single LLM Agent

Yinyi Luo, Yiqiao Jin, Weichen Yu +6

While large language model (LLM) multi-agent systems achieve superior reasoning performance through iterative debate, practical deployment is limited by their high computational co…

cs.SI2024

A Survey of Graph Neural Networks for Social Recommender Systems

Kartik Sharma, Yeon-Chang Lee, Sivagami Nambi +4

Social recommender systems (SocialRS) simultaneously leverage the user-to-item interactions as well as the user-to-user social relations for the task of generating item recommendat…

cs.CL2023

Findings of Factify 2: Multimodal Fake News Detection

S Suryavardan, Shreyash Mishra, Megha Chakraborty +9

With social media usage growing exponentially in the past few years, fake news has also become extremely prevalent. The detrimental impact of fake news emphasizes the need for rese…

cs.CL2025

A Thousand Words or An Image: Studying the Influence of Persona Modality in Multimodal LLMs

Julius Broomfield, Kartik Sharma, Srijan Kumar

Large language models (LLMs) have recently demonstrated remarkable advancements in embodying diverse personas, enhancing their effectiveness as conversational agents and virtual as…

cs.IR2024

SVD-AE: Simple Autoencoders for Collaborative Filtering

Seoyoung Hong, Jeongwhan Choi, Yeon-Chang Lee +2

Collaborative filtering (CF) methods for recommendation systems have been extensively researched, ranging from matrix factorization and autoencoder-based to graph filtering-based m…

hep-th2025

Large Wess-Zumino model at finite temperature and large chemical potential in

Srijan Kumar

We consider the supersymmetric Wess-Zumino model at large in dimension. We introduce a chemical potential() at finite temperature(). The non-trivial fixed point…

cs.CL2023

Memotion 3: Dataset on Sentiment and Emotion Analysis of Codemixed Hindi-English Memes

Shreyash Mishra, S Suryavardan, Parth Patwa +9

Memes are the new-age conveyance mechanism for humor on social media sites. Memes often include an image and some text. Memes can be used to promote disinformation or hatred, thus…

cs.SI2021

Evaluating Graph Vulnerability and Robustness using TIGER

Scott Freitas, Diyi Yang, Srijan Kumar +2

Network robustness plays a crucial role in our understanding of complex interconnected systems such as transportation, communication, and computer networks. While significant resea…