◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Nikolas Gritsch

5 papers hereh-index 486 citations7 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author2

Across the 4 of 5 papers where every author was matched, so the position is known.

fields
  • cs.CL3
  • cs.CY1
  • cs.LG1

identity via Semantic Scholar / OpenAlex

activity
20222025
most citedBAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts

1 citations · 1 across the 4 of their papers we have counts for

collaborators
Showing cs.CLShow all

3 papers · 1 filter

cs.CL2025

Command A: An Enterprise-Ready Large Language Model

Team Cohere, :, Aakanksha +227

In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command A is an agent-optimised…

cs.CL2024

Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts

Nikolas Gritsch, Qizhen Zhang, Acyr Locatelli +2

Efficiency, specialization, and adaptability to new data distributions are qualities that are hard to combine in current Large Language Models. The Mixture of Experts (MoE) archite…

cs.CL2023

Divergent Token Metrics: Measuring degradation to prune away LLM components -- and optimize quantization

Björn Deiseroth, Max Meuer, Nikolas Gritsch +4

Large Language Models (LLMs) have reshaped natural language processing with their impressive capabilities. However, their ever-increasing size has raised concerns about their effec…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.