27 citations · 44 across the 6 of their papers we have counts for
6 papers
Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
NVIDIA, :, Aaron Blakeman +571
We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 t…
SHERLOC: Structured Diagnostic Localization for Code Repair Agents
Hovhannes Tamoyan, Sean Narenthiran, Erik Arakelyan +2
LLM agents solve repository-level coding tasks through multi-turn tool use, but utilize half their budget on locating faults before editing. Dedicated localization frameworks have…
Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
Hovhannes Tamoyan, Subhabrata Dutta, Iryna Gurevych
Factual incorrectness in generated content is one of the primary concerns in ubiquitous deployment of large language models (LLMs). Prior findings suggest LLMs can (sometimes) dete…
LLM Roleplay: Simulating Human-Chatbot Interaction
Hovhannes Tamoyan, Hendrik Schuff, Iryna Gurevych
The development of chatbots requires collecting a large number of human-chatbot dialogues to reflect the breadth of users' sociodemographic backgrounds and conversational goals. Ho…
Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning
Lili Yu, Bowen Shi, Ramakanth Pasunuru +24
We present CM3Leon (pronounced "Chameleon"), a retrieval-augmented, token-based, decoder-only multi-modal language model capable of generating and infilling both text and images. C…
BARTSmiles: Generative Masked Language Models for Molecular Representations
Gayane Chilingaryan, Hovhannes Tamoyan, Ani Tevosyan +6
We discover a robust self-supervised strategy tailored towards molecular representations for generative masked language models through a series of tailored, in-depth ablations. Usi…