papers

Publications (6)

cs.LG2025

Apriel-Nemotron-15B-Thinker

Shruthan Radhakrishna, Soham Parikh, Gopal Sarda +32

While large language models (LLMs) have achieved remarkable reasoning capabilities across domains like code, math and other enterprise tasks, their significant memory and computati…

cs.CL2021

Predictive Representation Learning for Language Modeling

Qingfeng Lan, Luke Kumar, Martha White +1

To effectively perform the task of next-word prediction, long short-term memory networks (LSTMs) must keep track of many types of information. Some information is directly related…

cs.CL2018

On Generality and Knowledge Transferability in Cross-Domain Duplicate Question Detection for Heterogeneous Community Question Answering

Mohomed Shazan Mohomed Jabbar, Luke Kumar, Hamman Samuel +4

Duplicate question detection is an ongoing challenge in community question answering because semantically equivalent questions can have significantly different words and structures…

cs.LG2019

Gene Expression based Survival Prediction for Cancer Patients: A Topic Modeling Approach

Luke Kumar, Russell Greiner

Cancer is one of the leading cause of death, worldwide. Many believe that genomic data will enable us to better predict the survival time of these patients, which will lead to bett…

cs.LG2025

Apriel-H1: Towards Efficient Enterprise Reasoning Models

Oleksiy Ostapenko, Luke Kumar, Raymond Li +10

Large Language Models (LLMs) achieve remarkable reasoning capabilities through transformer architectures with attention mechanisms. However, transformers suffer from quadratic time…

cs.LG2025

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training

Oleksiy Ostapenko, Charles Guille-Escuret, Luke Kumar +7

We introduce a framework for optimizing domain-specific dataset construction in foundation model training. Specifically, we seek a cost-efficient way to estimate the quality of dat…