24 citations · 58 across the 50 of their papers we have counts for
27 papers · 1 filter
Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
Mobina Kashaniyan, Ali Jannesari
Test-time scaling can improve large language model reasoning by generating and combining multiple candidate responses. In sampling-based methods, the inference budget is often desc…
Multi-turn RL with Structural and Performance Aware Rewards for CUDA Kernel Generation
Quazi Ishtiaque Mahmud, Nesreen K. Ahmed, Ali Jannesari
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful technique to enhance the reasoning capacity of LLMs for optimized code generation. However, existing…
Sparse Subspace-to-Expert Sharing for Task-Agnostic Continual Learning
Fatema Siddika, Md Anwar Hossen, Tanwi Mallick +1
Continual learning in Large Language Models (LLMs) is hindered by the plasticity-stability dilemma, where acquiring new capabilities often leads to catastrophic forgetting of previ…
Fast MoE Inference via Predictive Prefetching and Expert Replication
Ankit Jyothish, Ali Jannesari, Aishwarya Sarkar +1
The Mixture of Experts (MoE) architecture has become a fundamental building block in state-of-the-art large language models (LLMs), improving domain-specific expertise in LLMs and…
CRAFT: Forgetting-Aware Intervention-Based Adaptation for Continual Learning
Md Anwar Hossen, Fatema Siddika, Juan Pablo Munoz +2
Large language models (LLMs) can acquire new capabilities through fine-tuning, but continual adaptation often leads to catastrophic forgetting. We propose CRAFT, a continual learni…
NOMAD: Generating Embeddings for Massive Distributed Graphs
Aishwarya Sarkar, Sayan Ghosh, Nathan R. Tallent +1
Successful machine learning on graphs or networks requires embeddings that not only represent nodes and edges as low-dimensional vectors but also preserve the graph structure. Esta…