5 papers
ProtoSSL: Interpretable Prototype Learning from Unlabeled Time-Series Data
Steven Song, Sahil Sethi, Brett Beaulieu-Jones +1
In time-series domains where both predictive performance and interpretability are essential, deep neural networks achieve strong results but provide limited insight into how their…
Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining
Anirudh Subramanyam, Yuxin Chen, Robert L. Grossman
Scaling laws for language model training traditionally characterize how performance scales with model size and dataset volume. Prior work has explored architecture variants and dat…
Multimodal Cancer Modeling in the Age of Foundation Model Embeddings
Steven Song, Morgan Borjigin-Wang, Irene Madejski +1
The Cancer Genome Atlas (TCGA) has enabled novel discoveries and served as a large-scale reference dataset in cancer through its harmonized genomics, clinical, and imaging data. Nu…
LaB-RAG: Label Boosted Retrieval Augmented Generation for Radiology Report Generation
Steven Song, Anirudh Subramanyam, Irene Madejski +1
In the current paradigm of image captioning, deep learning models are trained to generate text from image embeddings of latent features. We challenge the assumption that fine-tunin…
GDC Cohort Copilot: An AI Copilot for Curating Cohorts from the Genomic Data Commons
Steven Song, Anirudh Subramanyam, Zhenyu Zhang +2
The Genomic Data Commons (GDC) provides access to high quality, harmonized cancer genomics data through a unified curation and analysis platform centered around patient cohorts. Wh…