◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Nilabhra Roy Chowdhury

4 papers hereh-index 444 citations6 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author4

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CL2
  • cs.LG2

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.LG2025

ZClip: Adaptive Spike Mitigation for LLM Pre-Training

Abhay Kumar, Louis Owen, Nilabhra Roy Chowdhury +1

Training large language models (LLMs) presents numerous challenges, including gradient instability and loss spikes. These phenomena can lead to catastrophic divergence, requiring c…

cs.CL2025

A Refined Analysis of Massive Activations in LLMs

Louis Owen, Nilabhra Roy Chowdhury, Abhay Kumar +1

Motivated in part by their relevance for low-precision training and quantization, massive activations in large language models (LLMs) have recently emerged as a topic of interest.…

cs.LG2025

Variance Control via Weight Rescaling in LLM Pre-training

Louis Owen, Abhay Kumar, Nilabhra Roy Chowdhury +1

The outcome of Large Language Model (LLM) pre-training strongly depends on weight initialization and variance control strategies. Although the importance of initial variance contro…

cs.CL2024

Falcon2-11B Technical Report

Quentin Malartic, Nilabhra Roy Chowdhury, Ruxandra Cojocaru +14

We introduce Falcon2-11B, a foundation model trained on over five trillion tokens, and its multimodal counterpart, Falcon2-11B-vlm, which is a vision-to-text model. We report our f…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.