8 papers
Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation Models
Weimin Wu, Xuefeng Song, Yibo Wen +5
We introduce Genome-Factory, the first integrated Python library for tuning, deploying, and interpreting genomic foundation models. Our core contribution is to simplify and unify t…
Discrete Flow Matching Policy Optimization
Maojiang Su, Po-Chung Hsieh, Weimin Wu +4
We introduce Discrete flow Matching policy Optimization (DoMinO), a unified framework for Reinforcement Learning (RL) fine-tuning Discrete Flow Matching (DFM) models under a broad…
Sci2Pol: Evaluating and Fine-tuning LLMs on Scientific-to-Policy Brief Generation
Weimin Wu, Alexander C. Furnas, Eddie Yang +5
We propose Sci2Pol-Bench and Sci2Pol-Corpus, the first benchmark and training dataset for evaluating and fine-tuning large language models (LLMs) on policy brief generation from a…
Cell-JEPA: Latent Representation Learning for Single-Cell Transcriptomics
Ali ElSheikh, Rui-Xi Wang, Weimin Wu +9
Single-cell foundation models learn by reconstructing masked gene expression, implicitly treating technical noise as signal. With dropout rates exceeding 90%, reconstruction object…
On Structured State-Space Duality
Jerry Yao-Chieh Hu, Xiwen Zhang, Ali ElSheikh +2
Structured State-Space Duality (SSD) [Dao & Gu, ICML 2024] is an equivalence between a simple Structured State-Space Model (SSM) and a masked attention mechanism. In particular, a…
Universal Approximation with Softmax Attention
Jerry Yao-Chieh Hu, Hude Liu, Hong-Yu Chen +2
We prove that with linear transformations, both (i) two-layer self-attention and (ii) one-layer self-attention followed by a softmax function are universal approximators for contin…