activity
20242026
collaborators

6 papers

cs.CE2026

Cell-JEPA: Latent Representation Learning for Single-Cell Transcriptomics

Ali ElSheikh, Rui-Xi Wang, Weimin Wu +9

Single-cell foundation models learn by reconstructing masked gene expression, implicitly treating technical noise as signal. With dropout rates exceeding 90%, reconstruction object…

cs.LG2025

On Structured State-Space Duality

Jerry Yao-Chieh Hu, Xiwen Zhang, Ali ElSheikh +2

Structured State-Space Duality (SSD) [Dao & Gu, ICML 2024] is an equivalence between a simple Structured State-Space Model (SSM) and a masked attention mechanism. In particular, a…

cs.LG2025

POLO: Preference-Guided Multi-Turn Reinforcement Learning for Lead Optimization

Ziqing Wang, Yibo Wen, William Pattie +6

Lead optimization in drug discovery requires efficiently navigating vast chemical space through iterative cycles to enhance molecular properties while preserving structural similar…

cs.LG2025

Universal Approximation with Softmax Attention

Jerry Yao-Chieh Hu, Hude Liu, Hong-Yu Chen +2

We prove that with linear transformations, both (i) two-layer self-attention and (ii) one-layer self-attention followed by a softmax function are universal approximators for contin…

stat.ML2024

On Statistical Rates of Conditional Diffusion Transformers: Approximation, Estimation and Minimax Optimality

Jerry Yao-Chieh Hu, Weimin Wu, Yi-Chen Lee +3

We investigate the approximation and estimation rates of conditional diffusion transformers (DiTs) with classifier-free guidance. We present a comprehensive analysis for ``in-conte…

cs.LG2024

In-Context Deep Learning via Transformer Models

Weimin Wu, Maojiang Su, Jerry Yao-Chieh Hu +2

We investigate the transformer's capability to simulate the training process of deep models via in-context learning (ICL), i.e., in-context deep learning. Our key contribution is p…