21 papers
Small Models Scout Bottleneck Order for Large-Model Data Control
Seungmin Choi, Jiwon Sung, Muhammad Umer +4
Small proxy models are commonly used to identify data mixtures for larger-scale training. We ask whether their training trajectories reveal another transferable structure: the orde…
Beam-Response Contrastive Learning for Transmitter-Side MIMO CSI Representation
Sehyun Ryu, Yumin Kim, Minjae Lee +2
Self-supervised representation learning from unlabeled channel state information (CSI) can reduce labeling and adaptation overhead in learning-based multiple-input multiple-output…
General Preference Reinforcement Learning
Muhammad Umer, Muhammad Ahmed Mohsin, Ahsan Bilal +5
Post-training has split large language model (LLM) alignment into two largely disconnected tracks. Online reinforcement learning (RL) with verifiable rewards drives emergent reason…
Epistemic Uncertainty for Test-Time Discovery
Kainat Riaz, Muhammad Ahmed Mohsin, Ahsan Bilal +5
Automated scientific discovery using large language models relies on identifying genuinely novel solutions. Standard reinforcement learning penalizes high-variance mutations, which…
Continuous-Utility Direct Preference Optimization
Muhammad Ahmed Mohsin, Muhammad Umer, Ahsan Bilal +6
Large language model reasoning is often treated as a monolithic capability, relying on binary preference supervision that fails to capture partial progress or fine-grained reasonin…
Neural Gaussian Radio Fields for Channel Estimation
Muhammad Umer, Muhammad Ahmed Mohsin, Ahsan Bilal +1
Accurate channel state information (CSI) is a critical bottleneck in modern wireless networks, with pilot overhead consuming 11\% to 21\% of transmission bandwidth and feedback del…