papers

Publications (34)

quant-ph2025

Efficient Privacy-Preserving Training of Quantum Neural Networks by Using Mixed States to Represent Input Data Ensembles

Gaoyuan Wang, Jonathan Warrell, Mark Gerstein

Quantum neural networks (QNNs) are gaining increasing interest due to their potential to detect complex patterns in data by leveraging uniquely quantum phenomena. This makes them p…

cs.CL2024

Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data?

Xiangru Tang, Yiming Zong, Jason Phang +4

Despite the remarkable capabilities of Large Language Models (LLMs) like GPT-4, producing complex, structured tabular data remains challenging. Our study assesses LLMs' proficiency…

cs.LG2026

MIOFlow 2.0: A unified framework for inferring cellular stochastic dynamics from single cell and spatial transcriptomics data

Xingzhi Sun, João Felipe Rocha, Brett Phelan +11

Understanding cellular trajectories via time-resolved single-cell transcriptomics is vital for studying development, regeneration, and disease. A key challenge is inferring continu…

q-bio.BM2026

SurfDesign: Effective Protein Design on Molecular Surfaces

Fang Wu, Shuting Jin, Xiangru Tang +5

Protein function is largely determined by molecular surface geometry and physicochemical complementarity, yet most protein design methods condition only on backbone structure. We i…

cs.CL2023

GersteinLab at MEDIQA-Chat 2023: Clinical Note Summarization from Doctor-Patient Conversations through Fine-tuning and In-context Learning

Xiangru Tang, Andrew Tran, Jeffrey Tan +1

This paper presents our contribution to the MEDIQA-2023 Dialogue2Note shared task, encompassing both subtask A and subtask B. We approach the task as a dialogue summarization probl…

cs.LG2026

MPP-GNN: Subject-Adaptive Community Detection for fMRI-Based Alzheimer's Disease Classification

Yang Zhang, Xiao Zhou, Jonathan Warrell +3

Functional magnetic resonance imaging (fMRI) is a widely used technique for studying the brain. Recent methods that utilize graph neural networks (GNNs) for analysis of brain funct…

cs.CL2024

MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning

Xiangru Tang, Anni Zou, Zhuosheng Zhang +5

Large language models (LLMs), despite their remarkable progress across various general domains, encounter significant barriers in medicine and healthcare. This field faces unique c…

cs.CL2024

Investigating Data Contamination in Modern Benchmarks for Large Language Models

Chunyuan Deng, Yilun Zhao, Xiangru Tang +2

Recent observations have underscored a disparity between the inflated benchmark scores and the actual performance of LLMs, raising concerns about potential contamination of evaluat…

cs.CL2024

ChatCell: Facilitating Single-Cell Analysis with Natural Language

Yin Fang, Kangwei Liu, Ningyu Zhang +7

As Large Language Models (LLMs) rapidly evolve, their influence in science is becoming increasingly prominent. The emerging capabilities of LLMs in task generalization and free-for…

cs.LG2022

Forest Fire Clustering for Single-cell Sequencing with Iterative Label Propagation and Parallelized Monte Carlo Simulation

Zhanlin Chen, Jeremy Goldwasser, Philip Tuckman +3

In the era of single-cell sequencing, there is a growing need to extract insights from data with clustering methods. Here, we introduce Forest Fire Clustering, an efficient and int…

cs.CL2025

ChemAgent: Self-updating Library in Large Language Models Improves Chemical Reasoning

Xiangru Tang, Tianyu Hu, Muyang Ye +9

Chemical reasoning usually involves complex, multi-step processes that demand precise calculations, where even minor errors can lead to cascading failures. Furthermore, large langu…

cs.LG2026

CellForge: Agentic Design of Virtual Cell Models

Xiangru Tang, Zhuoyun Yu, Jiapeng Chen +12

Virtual cell modeling aims to predict cellular responses to diverse perturbations but faces challenges from biological complexity, multimodal data heterogeneity, and the need for i…

cs.CL2026

MedicalAgentsBench for Complex Medical Reasoning: Comparing Internalized Reasoning Models versus Externalized Agent-based Frameworks

Yanjun Shao, Xiangru Tang, Jiwoong Sohn +10

Complex medical reasoning requires integrating heterogeneous clinical evidence across multiple inference steps. Large language models (LLMs) now approach this through two routes: i…

cs.CL2024

ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain

Haochen Zhao, Xiangru Tang, Ziran Yang +8

The advancement and extensive application of large language models (LLMs) have been remarkable, including their use in scientific research assistance. However, these models often g…

cs.CY2025

Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy

Xiangru Tang, Qiao Jin, Kunlun Zhu +10

AI scientists powered by large language models have demonstrated substantial promise in autonomously conducting experiments and facilitating scientific discoveries across various d…

cs.LG2025

STAGED: A Multi-Agent Neural Network for Learning Cellular Interaction Dynamics

Joao F. Rocha, Ke Xu, Xingzhi Sun +6

The advent of single-cell technology has significantly improved our understanding of cellular states and subpopulations in various tissues under normal and diseased conditions by e…

cs.CL2023

Igniting Language Intelligence: The Hitchhiker's Guide From Chain-of-Thought Reasoning to Language Agents

Zhuosheng Zhang, Yao Yao, Aston Zhang +8

Large language models (LLMs) have dramatically enhanced the field of language intelligence, as demonstrably evidenced by their formidable empirical performance across a spectrum of…

cs.AI2023

ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Yujia Qin, Shihao Liang, Yining Ye +16

Despite the advancements of open-source large language models (LLMs), e.g., LLaMA, they remain significantly limited in tool-use capabilities, i.e., using external tools (APIs) to…

cs.AI2026

Advancing AI Research Assistants with Expert-Involved Learning

Tianyu Liu, Simeng Han, Hanchen Wang +27

Large language models (LLMs) and large multimodal models (LMMs) promise to accelerate biomedical discovery, yet their reliability remains unclear. We introduce ARIEL (AI Research A…

cs.LG2025

E2Former: An Efficient and Equivariant Transformer with Linear-Scaling Tensor Products

Yunyang Li, Lin Huang, Zhihao Ding +10

Equivariant Graph Neural Networks (EGNNs) have demonstrated significant success in modeling microscale systems, including those in chemistry, biology and materials science. However…

cs.CL2024

MIMIR: A Streamlined Platform for Personalized Agent Tuning in Domain Expertise

Chunyuan Deng, Xiangru Tang, Yilun Zhao +5

Recently, large language models (LLMs) have evolved into interactive agents, proficient in planning, tool use, and task execution across a wide variety of tasks. However, without s…

cs.LG2024

BioCoder: A Benchmark for Bioinformatics Code Generation with Large Language Models

Xiangru Tang, Bill Qian, Rick Gao +3

Pre-trained large language models (LLMs) have significantly improved code generation. As these models scale up, there is an increasing need for the output to handle more intricate…

cs.CR2022

Scalable privacy-preserving cancer type prediction with homomorphic encryption

Esha Sarkar, Eduardo Chielle, Gamze Gursoy +3

Machine Learning (ML) alleviates the challenges of high-dimensional data analysis and improves decision making in critical applications like healthcare. Effective cancer type from…

cs.CL2025

Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards

Jaehoon Yun, Jiwoong Sohn, Jungwoo Park +9

Large language models have shown promise in clinical decision making, but current approaches struggle to localize and correct errors at specific steps of the reasoning process. Thi…

cs.CE2026

D-Flow: Multi-modality Flow Matching for D-peptide Design

Fang Wu, Shuting Jin, Xiangru Tang +3

Among these, D-peptides are resistant to proteolysis, exhibit greater in vivo stability, and are easier to synthesize. Despite advances in deep learning for peptide discovery, the…

cs.LG2026

Elign: Equivariant Diffusion Model Alignment from Foundational Machine Learning Force Fields

Yunyang Li, Lin Huang, Luojia Xia +2

Generative models for 3D molecular conformations must respect Euclidean symmetries and concentrate probability mass on thermodynamically favorable, mechanically stable structures.…

q-bio.BM2024

A Survey of Generative AI for de novo Drug Design: New Frontiers in Molecule and Protein Generation

Xiangru Tang, Howard Dai, Elizabeth Knight +4

Artificial intelligence (AI)-driven methods can vastly improve the historically costly drug design process, with various generative models already in widespread use. Generative mod…

cs.LG2022

Higher-Order Generalization Bounds: Learning Deep Probabilistic Programs via PAC-Bayes Objectives

Jonathan Warrell, Mark Gerstein

Deep Probabilistic Programming (DPP) allows powerful models based on recursive computation to be learned using efficient deep-learning optimization techniques. Additionally, DPP of…

cs.CL2024

Step-Back Profiling: Distilling User History for Personalized Scientific Writing

Xiangru Tang, Xingyao Zhang, Yanjun Shao +6

Large language models (LLM) excel at a variety of natural language processing tasks, yet they struggle to generate personalized content for individuals, particularly in real-world…

cs.CL2024

ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Xiangru Tang, Yuliang Liu, Zefan Cai +21

Despite Large Language Models (LLMs) like GPT-4 achieving impressive results in function-level code generation, they struggle with repository-scale code understanding (e.g., coming…

quant-ph2025

-QVAE: A Quantum Variational Autoencoder utilizing Regularized Mixed-state Latent Representations

Gaoyuan Wang, Jonathan Warrell, Prashant S. Emani +1

A major challenge in quantum computing is its application to large real-world datasets due to scarce quantum hardware resources. One approach to enabling tractable quantum models f…

q-bio.BM2023

Disentangled Wasserstein Autoencoder for T-Cell Receptor Engineering

Tianxiao Li, Hongyu Guo, Filippo Grazioli +2

In protein biophysics, the separation between the functionally important residues (forming the active site or binding surface) and those that create the overall structure (the fold…

cs.LG2026

Dispersion Loss Counteracts Embedding Condensation and Improves Generalization in Small Language Models

Chen Liu, Xingzhi Sun, Xi Xiao +8

Large language models (LLMs) achieve remarkable performance through ever-increasing parameter counts, but scaling incurs steep computational costs. To better understand LLM scaling…

cs.LG2018

Rank Projection Trees for Multilevel Neural Network Interpretation

Jonathan Warrell, Hussein Mohsen, Mark Gerstein

A variety of methods have been proposed for interpreting nodes in deep neural networks, which typically involve scoring nodes at lower layers with respect to their effects on the o…