papers

Publications (31)

stat.ME2026

Dynamic Frechet Regression with Feature Selection for Distributional Data

Kiran Adhikari, Amrutha Dinesh, Mathew Kuttolamadom +1

Many scientific and engineering applications generate responses that are not scalars or vectors, but statistical objects whose form evolves over an ordered index such as time, dept…

cs.CL2025

NVIDIA Nemotron 3: Efficient and Open Intelligence

NVIDIA, :, Aaron Blakeman +356

We introduce the Nemotron 3 family of models - Nano, Super, and Ultra. These models deliver strong agentic, reasoning, and conversational capabilities. The Nemotron 3 family uses a…

cs.CL2026

Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

NVIDIA, :, Aaron Blakeman +571

We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 t…

cs.LG2026

Beyond Static Bias: Adaptive Multi-Fidelity Bandits with Improving Proxies

Muyun Lu, Haoyang Hong, Huazheng Wang +1

As an extension of the classical multi-armed bandit problem, multi-fidelity multi-armed bandits (MF-MAB) enable individual arms to be evaluated using diverse feedback sources that…

cs.LG2026

Nemotron-CrossThink: Scaling Self-Learning beyond Math Reasoning

Syeda Nahida Akter, Shrimai Prabhumoye, Matvei Novikov +8

Large Language Models (LLMs) have shown strong reasoning capabilities, particularly when enhanced through Reinforcement Learning (RL). While prior work has successfully applied RL…

cs.CL2017

Acquiring Background Knowledge to Improve Moral Value Prediction

Ying Lin, Joe Hoover, Morteza Dehghani +2

In this paper, we address the problem of detecting expressions of moral values in tweets using content analysis. This is a particularly challenging problem because moral values are…

cs.AI2026

Are Tools All We Need? Unveiling the Tool-Use Tax in LLM Agents

Kaituo Zhang, Zhen Xiong, Mingyu Zhong +4

Tool-augmented reasoning has become a popular direction for LLM-based agents, and it is widely assumed to improve reasoning and reliability. However, we demonstrate that this conse…

cs.CR2016

GID: Graph-based Intrusion Detection on Massive Process Traces for Enterprise Security Systems

Boxiang Dong, Zhengzhang Chen, Hui Wang +5

Intrusion detection system (IDS) is an important part of enterprise security system architecture. In particular, anomaly-based IDS has been widely applied to detect abnormal proces…

cs.CL2021

Personalized Entity Resolution with Dynamic Heterogeneous Knowledge Graph Representations

Ying Lin, Han Wang, Jiangning Chen +5

The growing popularity of Virtual Assistants poses new challenges for Entity Resolution, the task of linking mentions in text to their referent entities in a knowledge base. Specif…

math.OC2024

Tight error bounds for log-determinant cones without constraint qualifications

Ying Lin, Scott B. Lindstrom, Bruno F. Lourenço +1

In this paper, without requiring any constraint qualifications, we establish tight error bounds for the log-determinant cone, which is the closure of the hypograph of the perspecti…

cs.CR2018

Collaborative Alerts Ranking for Anomaly Detection

Ying Lin, Zhengzhang Chen, Cheng Cao +5

Given a large number of low-level heterogeneous categorical alerts from an anomaly detection system, how to characterize complex relationships between different alerts, filter out…

math.OC2023

Generalized power cones: optimal error bounds and automorphisms

Ying Lin, Scott B. Lindstrom, Bruno F. Lourenço +1

Error bounds are a requisite for trusting or distrusting solutions in an informed way. Until recently, provable error bounds in the absence of constraint qualifications were unatta…

cs.CL2025

NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model

NVIDIA, :, Aarti Basant +214

We introduce Nemotron-Nano-9B-v2, a hybrid Mamba-Transformer language model designed to increase throughput for reasoning workloads while achieving state-of-the-art accuracy compar…

cs.CL2025

Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset

Dan Su, Kezhi Kong, Ying Lin +6

Recent English Common Crawl datasets like FineWeb-Edu and DCLM achieved significant benchmark gains via aggressive model-based filtering, but at the cost of removing 90% of data. T…

cs.CL2019

A Grounded Unsupervised Universal Part-of-Speech Tagger for Low-Resource Languages

Ronald Cardenas, Ying Lin, Heng Ji +1

Unsupervised part of speech (POS) tagging is often framed as a clustering problem, but practical taggers need to \textit{ground} their clusters as well. Grounding generally require…

cs.LG2026

Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

NVIDIA, :, Aakshita Chandiramani +544

We describe the pre-training, post-training, and quantization of Nemotron 3 Super, a 120 billion (active 12 billion) parameter hybrid Mamba-Attention Mixture-of-Experts model. Nemo…

cs.AI2026

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data

Kaituo Zhang, Mingzhi Hu, Hoang Anh Duy Le +9

Large Language Models (LLMs) have emerged as powerful tools for generating data across various modalities. By transforming data from a scarce resource into a controllable asset, LL…

stat.ML2024

Ranking and Combining Latent Structured Predictive Scores without Labeled Data

Shiva Afshar, Yinghan Chen, Shizhong Han +1

Combining multiple predictors obtained from distributed data sources to an accurate meta-learner is promising to achieve enhanced performance in lots of prediction problems. As the…

cs.CL2025

Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

NVIDIA, :, Aaron Blakeman +311

We present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more than 3 t…

cs.LG2021

Adaptive perturbation adversarial training: based on reinforcement learning

Zhishen Nie, Ying Lin, Sp Ren +1

Adversarial training has become the primary method to defend against adversarial samples. However, it is hard to practically apply due to many shortcomings. One of the shortcomings…

stat.ME2026

Change-point detection in variance-covariance matrix

Ying Lin, Benjamin Poignard

We consider the joint estimation of change point locations and the sparsity pattern of the variance covariance matrix, which is assumed to evolve in a piecewise constant manner. By…

cs.CL2025

Llama-Nemotron: Efficient Reasoning Models

Akhiad Bercovich, Itay Levy, Izik Golan +132

We introduce the Llama-Nemotron series of models, an open family of heterogeneous reasoning models that deliver exceptional reasoning capabilities, inference efficiency, and an ope…

math.ST2026

Change Point Detection in Precision Matrices with D-trace Loss

Ying Lin, Benjamin Poignard, Ting Kei Pong +1

We consider the problem of estimating a time-varying sparse precision matrix, which is assumed to evolve in a piecewise constant manner. Building upon the Group Fused LASSO and LAS…

cs.LG2024

FCOM: A Federated Collaborative Online Monitoring Framework via Representation Learning

Tanapol Kosolwattana, Huazheng Wang, Raed Al Kontar +1

Online learning has demonstrated notable potential to dynamically allocate limited resources to monitor a large population of processes, effectively balancing the exploitation of p…

stat.ML2022

A Generative Adversarial Network-based Selective Ensemble Characteristic-to-Expression Synthesis (SE-CTES) Approach and Its Applications in Healthcare

Yuxuan Li, Ying Lin, Chenang Liu

Investigating the causal relationships between characteristics and expressions plays a critical role in healthcare analytics. Effective synthesis for expressions using given charac…

cs.LG2023

Online Modeling and Monitoring of Dependent Processes under Resource Constraints

Tanapol Kosolwattana, Huazheng Wang, Ying Lin

Adaptive monitoring of a large population of dynamic processes is critical for the timely detection of abnormal events under limited resources in many healthcare and engineering sy…

q-bio.NC2020

Cost-efficiency trade-offs of the human brain network revealed by a multiobjective evolutionary algorithm

Junji Ma, Jinbo Zhang, Ying Lin +1

It is widely believed that the formation of brain network structure is under the pressure of optimal trade-off between reducing wiring cost and promoting communication efficiency.…

cs.CL2021

COVID-19 Literature Knowledge Graph Construction and Drug Repurposing Report Generation

Qingyun Wang, Manling Li, Xuan Wang +24

To combat COVID-19, both clinicians and scientists need to digest vast amounts of relevant biomedical knowledge in scientific literature to understand the disease mechanism and rel…

cs.LG2025

Decentralized Optimization with Topology-Independent Communication

Ying Lin, Yao Kuang, Ahmet Alacaoglu +1

Distributed optimization requires nodes to coordinate, yet full synchronization scales poorly. When nodes collaborate through pairwise regularizers, standard methods demand…

cs.CL2025

Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

NVIDIA, :, Aaron Blakeman +198

As inference-time scaling becomes critical for enhanced reasoning capabilities, it is increasingly becoming important to build models that are efficient to infer. We introduce Nemo…

stat.ML2026

Permutation-preserving Functions and Neural Vecchia Covariance Kernels

Jian Cao, Nian Liu, Ying Lin

We introduce a novel framework for constructing scalable and flexible covariance kernels for Gaussian processes (GPs) by directly learning the covariance structure under a regressi…