◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Christoph Schuhmann

9 papers hereh-index 59.4k citations12 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author5
  • middle author2

Across the 7 of 9 papers where every author was matched, so the position is known.

fields
  • cs.CV4
  • cs.CL3
  • cs.LG2

identity via Semantic Scholar / OpenAlex

activity
20212026
most citedLAION-5B: An open large-scale dataset for training next generation image-text models

1k citations · 1.4k across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

3 papers · 1 filter

cs.CL2025

MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources

Huu Nguyen, Victor May, Harsh Raj +14

We present MixtureVitae, an open-access pretraining corpus built to minimize legal risk while providing strong downstream performance. MixtureVitae follows a permissive-first, risk…

cs.CL2025

EmoNet-Voice: A Fine-Grained, Expert-Verified Benchmark for Speech Emotion Detection

Christoph Schuhmann, Robert Kaczmarczyk, Gollam Rabby +6

Speech emotion recognition (SER) systems are constrained by existing datasets that typically cover only 6-10 basic emotions, lack scale and diversity, and face ethical challenges w…

cs.CL2023

OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Andreas Köpf, Yannic Kilcher, Dimitri von Rütte +15

Aligning large language models (LLMs) with human preferences has proven to drastically improve usability and has driven rapid adoption as demonstrated by ChatGPT. Alignment techniq…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.