◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Taiqi He

4 papers hereh-index 8135 citations10 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author2

Across the 3 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CL4
ORCID 0000-0002-7122-3493

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.CL2024

GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text

Michael Ginn, Lindia Tjuatja, Taiqi He +4

Language documentation projects often involve the creation of annotated text in a format such as interlinear glossed text (IGT), which captures fine-grained morphosyntactic analyse…

cs.CL2024

Hire a Linguist!: Learning Endangered Languages with In-Context Linguistic Descriptions

Kexun Zhang, Yee Man Choi, Zhenqiao Song +3

How can large language models (LLMs) process and translate endangered languages? Many languages lack a large corpus to train a decent LLM; therefore existing LLMs rarely perform we…

cs.CL2024

Wav2Gloss: Generating Interlinear Glossed Text from Speech

Taiqi He, Kwanghee Choi, Lindia Tjuatja +6

Thousands of the world's languages are in danger of extinction--a tremendous threat to cultural identities and human language diversity. Interlinear Glossed Text (IGT) is a form of…

cs.CL2024

Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons

Shijia Zhou, Leonie Weissweiler, Taiqi He +3

In this paper, we make a contribution that can be understood from two perspectives: from an NLP perspective, we introduce a small challenge dataset for NLI with large lexical overl…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.