◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Thuat Nguyen-Tran

2 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author1

Across the 2 of 2 papers where every author was matched, so the position is known.

fields
  • cs.CL2
ORCID 0000-0002-3761-2794

identity via Semantic Scholar / OpenAlex

most citedCulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

19 citations · 22 across the 2 of their papers we have counts for

collaborators

2 papers

cs.CL2023★ 19 cited

CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai +5

The driving factors behind the development of large language models (LLMs) with impressive learning capabilities are their colossal model sizes and extensive training datasets. Alo…

cs.CL2023★ 3 cited

Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback

Viet Dac Lai, Chien Van Nguyen, Nghia Trung Ngo +4

A key technology for the development of large language models (LLMs) involves instruction tuning that helps align the models' responses with human expectations to realize impressiv…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.