21 papers
Auditing pretraining contamination in single-cell foundation model benchmarks
Sarwan Ali
Single-cell foundation models (scFMs) such as Geneformer, scGPT, and Universal Cell Embeddings (UCE) are pretrained on tens of millions of cells drawn from public repositories. The…
Causal dictionary learning reveals and validates transcription-factor binding features in genomic language models
Sarwan Ali
Genomic language models achieve strong performance across regulatory-genomics tasks, yet what these models internally represent remains opaque, and the field lacks a principled pro…
Multi-Scale Reversible Chaos Game Representation: A Unified Framework for Sequence Classification
Sarwan Ali, Taslim Murad
Biological classification with interpretability remains a challenging task. For this, we introduce a novel encoding framework, Multi-Scale Reversible Chaos Game Representation (MS-…
Uncovering Hierarchical Structure in LLM Embeddings with -Hyperbolicity, Ultrametricity, and Neighbor Joining
Prakash Chourasia, Sarwan Ali, Murray Patterson
The rapid advancement of large language models (LLMs) has enabled significant strides in various fields. This paper introduces a novel approach to evaluate the effectiveness of LLM…
DPSR: Differentially Private Sparse Reconstruction via Multi-Stage Denoising for Recommender Systems
Sarwan Ali
Differential privacy (DP) has emerged as the gold standard for protecting user data in recommender systems, but existing privacy-preserving mechanisms face a fundamental challenge:…
Boosting t-SNE Efficiency for Sequencing Data: Insights from Kernel Selection
Avais Jan, Prakash Chourasia, Sarwan Ali +1
Dimensionality reduction techniques are essential for visualizing and analyzing high-dimensional biological sequencing data. t-distributed Stochastic Neighbor Embedding (t-SNE) is…