papers

Publications (31)

cs.AI2026

VAmoS Bench: Voice Agent Simulation Bench

Joshua Meyer, Sahar Shayegan, Ritiz Tambi +5

The paper presents VAmoS Bench, a simulation-based benchmark that evaluates complete voice‑agent systems on end‑to‑end customer‑support tasks, checking both conversational behavior…

#voice agents#benchmark#simulation#customer support
math.NT2022

Exact evaluations and reciprocity theorems for finite trigonometric sums

Bruce C. Berndt, Sun Kim, Alexandru Zaharescu

We evaluate in closed form several classes of finite trigonometric sums. Two general methods are used. The first is new and involves sums of roots of unity. The second uses contour…

math.NT2022

Finite trigonometric sums arising from Ramanujan's theta functions

Bruce C. Berndt, Sun Kim, Alexandru Zaharescu

Two classes of finite trigonometric sums, each involving only 's, are evaluated in closed form. The previous and original proofs arise from Ramanujan's theta functions and mo…

cs.LG2026

ISAAC: Auditing Causal Reasoning in Deep Models for Drug-Target Interaction

Barbara Tarantino, Sun Kim, Yijingxiu Lu +1

Deep learning models for drug--target interaction (DTI) prediction often achieve strong benchmark performance without necessarily relying on mechanistically meaningful molecular fe…

cs.HC2020

TeamTat: a collaborative text annotation tool

Rezarta Islamaj, Dongseop Kwon, Sun Kim +1

Manually annotated data is key to developing text-mining and information-extraction algorithms. However, human annotation requires considerable time, effort and expertise. Given th…

math.NT2017

Sums of squares and products of Bessel functions

Bruce C. Berndt, Atul Dixit, Sun Kim +1

Let denote the number of representations of the positive integer as the sum of squares. We rigorously prove for the first time a Voronoi summation formula for $r_k…

cs.LG2025

SubDyve: Subgraph-Driven Dynamic Propagation for Virtual Screening Enhancement Controlling False Positive

Jungseob Yi, Seoyoung Choi, Sun Kim +1

Virtual screening (VS) aims to identify bioactive compounds from vast chemical libraries, but remains difficult in low-label regimes where only a few actives are known. Existing me…

cs.CL2019

Deep learning with sentence embeddings pre-trained on biomedical corpora improves the performance of finding similar sentences in electronic medical records

Qingyu Chen, Jingcheng Du, Sun Kim +2

Capturing sentence semantics plays a vital role in a range of text mining applications. Despite continuous efforts on the development of related datasets and models in the general…

cs.CL2022

Sparse Structure Learning via Graph Neural Networks for Inductive Document Classification

Yinhua Piao, Sangseon Lee, Dohoon Lee +1

Recently, graph neural networks (GNNs) have been widely used for document classification. However, most existing methods are based on static word co-occurrence graphs without sente…

cs.LG2022

Triangular Contrastive Learning on Molecular Graphs

MinGyu Choi, Wonseok Shin, Yijingxiu Lu +1

Recent contrastive learning methods have shown to be effective in various tasks, learning generalizable representations invariant to data augmentation thereby leading to state of t…

cs.LG2024

Improving out-of-distribution generalization in graphs via hierarchical semantic environments

Yinhua Piao, Sangseon Lee, Yijingxiu Lu +1

Out-of-distribution (OOD) generalization in the graph domain is challenging due to complex distribution shifts and a lack of environmental contexts. Recent methods attempt to enhan…

cs.LG2026

CombiMOTS: Combinatorial Multi-Objective Tree Search for Dual-Target Molecule Generation

Thibaud Southiratn, Bonil Koo, Yijingxiu Lu +1

Dual-target molecule generation, which focuses on discovering compounds capable of interacting with two target proteins, has garnered significant attention due to its potential for…

math.NT2021

Two-parameter Identities for Divisor Sums in Algebraic Number Fields

Bruce C. Berndt, Martino Fassina, Sun Kim +1

In a one-page fragment published with his lost notebook, Ramanujan stated two double series identities associated, respectively, with the famous Gauss Circle and Dirichlet Divisor…

cs.CL2025

MV-CLAM: Multi-View Molecular Interpretation with Cross-Modal Projection via Language Model

Sumin Ha, Jun Hyeong Kim, Yinhua Piao +1

Human expertise in chemistry and biomedicine relies on contextual molecular understanding, a capability that large language models (LLMs) can extend through fine-grained alignment…

cs.CV2018

Hybrid Approach of Relation Network and Localized Graph Convolutional Filtering for Breast Cancer Subtype Classification

Sungmin Rhee, Seokjun Seo, Sun Kim

Network biology has been successfully used to help reveal complex mechanisms of disease, especially cancer. On the other hand, network biology requires in-depth knowledge to constr…

cs.CL2023

Clinical Note Owns its Hierarchy: Multi-Level Hypergraph Neural Networks for Patient-Level Representation Learning

Nayeon Kim, Yinhua Piao, Sun Kim

Leveraging knowledge from electronic health records (EHRs) to predict a patient's condition is essential to the effective delivery of appropriate care. Clinical notes of patient EH…

cs.LG2026

Predicting Therapeutic Outcome via Aligning Patient-Specific Knowledge Graph and Gene-Level Perturbation Representations

Dongmin Bang, Sugyun An, Inyoung Sung +3

Accurate prediction of patient-specific therapeutic response from pre-treatment transcriptomes is hindered by the scarcity of matched clinical response labels and post-treatment mo…

math.NT2020

The distribution of spacings between the fractional parts of

Martino Fassina, Sun Kim, Alexandru Zaharescu

We study the distribution of spacings between the fractional parts of . For of high enough Diophantine type we prove a necessary and sufficient condition for $n^dα\mod…

cs.LG2023

SPGP: Structure Prototype Guided Graph Pooling

Sangseon Lee, Dohoon Lee, Yinhua Piao +1

While graph neural networks (GNNs) have been successful for node classification tasks and link prediction tasks in graph, learning graph-level representations still remains a chall…

cs.CV2025

AI-driven Automation of End-to-end Assessment of Suturing Expertise

Atharva Deo, Nicholas Matsumoto, Sun Kim +8

We present an AI based approach to automate the End-to-end Assessment of Suturing Expertise (EASE), a suturing skills assessment tool that comprehensively defines criteria around r…

cs.IR2018

A Fast Deep Learning Model for Textual Relevance in Biomedical Information Retrieval

Sunil Mohan, Nicolas Fiorini, Sun Kim +1

Publications in the life sciences are characterized by a large technical vocabulary, with many lexical and semantic variations for expressing the same concept. Towards addressing t…

math.NT2016

On a theorem of A. I. Popov on sums of squares

Bruce C. Berndt, Atul Dixit, Sun Kim +1

Let denote the number of representations of the positive integer as the sum of squares. In 1934, the Russian mathematician A.~I.~Popov stated, but did not rigorous…

cs.CL2024

SLM as Guardian: Pioneering AI Safety with Small Language Models

Ohjoon Kwon, Donghyeon Jeon, Nayoung Choi +6

Most prior safety research of large language models (LLMs) has focused on enhancing the alignment of LLMs to better suit the safety requirements of humans. However, internalizing s…

math.NT2025

Ninth degree analogue of Ramanujan's septic theta function identity

Sun Kim, Örs Rebák

On page 206 in his lost notebook, Ramanujan recorded a seventh degree identity for his theta function . We give an analogous ninth degree identity. We also provide an applic…

cs.LG2025

GraphT5: Unified Molecular Graph-Language Modeling via Multi-Modal Cross-Token Attention

Sangyeup Kim, Nayeon Kim, Yinhua Piao +1

Molecular language modeling tasks such as molecule captioning have been recognized for their potential to further understand molecular properties that can aid drug discovery or mat…

cs.CL2017

Bridging the Gap: Incorporating a Semantic Similarity Measure for Effectively Mapping PubMed Queries to Documents

Sun Kim, Nicolas Fiorini, W. John Wilbur +1

The main approach of traditional information retrieval (IR) is to examine how many words from a query appear in a document. A drawback of this approach, however, is that it may fai…

cs.IR2025

Taxonomy and Analysis of Sensitive User Queries in Generative AI Search

Hwiyeol Jo, Taiwoo Park, Hyunwoo Lee +10

Although there has been a growing interest among industries in integrating generative LLMs into their services, limited experience and scarcity of resources act as a barrier in lau…

cs.LG2021

Handling Long-Tail Queries with Slice-Aware Conversational Systems

Cheng Wang, Sun Kim, Taiwoo Park +5

We have been witnessing the usefulness of conversational AI systems such as Siri and Alexa, directly impacting our daily lives. These systems normally rely on machine learning mode…

math.NT2021

Balanced Derivatives, Identities, and Bounds for Trigonometric and Bessel Series

Bruce C. Berndt, Martino Fassina, Sun Kim +1

Motivated by two identities published with Ramanujan's lost notebook and connected, respectively, with the Gauss circle problem and the Dirichlet divisor problem, in an earlier pap…

math.NT2024

Evaluations and relations for finite trigonometric sums

Bruce C. Berndt, Sun Kim, Alexandru Zaharescu

Several methods are used to evaluate finite trigonometric sums. In each case, either the sum had not previously been evaluated, or it had been evaluated, but only by analytic means…

cs.CL2019

BioConceptVec: creating and evaluating literature-based biomedical concept embeddings on a large scale

Qingyu Chen, Kyubum Lee, Shankai Yan +3

Capturing the semantics of related biological concepts, such as genes and mutations, is of significant importance to many research tasks in computational biology such as protein-pr…