4 papers
Small Models Scout Bottleneck Order for Large-Model Data Control
Seungmin Choi, Jiwon Sung, Muhammad Umer +4
Small proxy models are commonly used to identify data mixtures for larger-scale training. We ask whether their training trajectories reveal another transferable structure: the orde…
Epistemic Uncertainty for Test-Time Discovery
Kainat Riaz, Muhammad Ahmed Mohsin, Ahsan Bilal +5
Automated scientific discovery using large language models relies on identifying genuinely novel solutions. Standard reinforcement learning penalizes high-variance mutations, which…
: Stratified Scaling Search for Test-Time in Diffusion Language Models
Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer +6
Test-time scaling investigates whether a fixed diffusion language model (DLM) can generate better outputs when given more inference compute, without additional training. However, n…
Explainable AI in Genomics: Transcription Factor Binding Site Prediction with Mixture of Experts
Aakash Tripathi, Ian E. Nielsen, Muhammad Umer +2
Transcription Factor Binding Site (TFBS) prediction is crucial for understanding gene regulation and various biological processes. This study introduces a novel Mixture of Experts…