◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Chetan Phakami Pun

3 papers hereh-index 125 citations3 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • last author3

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL2
  • cs.LG1

identity via Semantic Scholar / OpenAlex

most citedGPT Semantic Cache: Reducing LLM Costs and Latency via Semantic Embedding Caching

1 citations · 1 across the 3 of their papers we have counts for

collaborators

3 papers

cs.CL2026

Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer

Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun

We present Arkios, a 1.04B-parameter dense transformer pretrained from scratch on 150B tokens of bilingual English-Nepali text, using a custom single-file C/CUDA training stack and…

cs.CL2026

Vowel Signs Are Not Letters: A Pre-tokenization Ceiling on Multilingual Tokenizer Fertility

Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun

Byte-level BPE tokenizers that use the HuggingFace ByteLevel pre-tokenizer inherit GPT-2's word regex, where a word is defined as \p{L}+, one or more Unicode letters. In abugida sc…

cs.LG2024★ 1 cited

GPT Semantic Cache: Reducing LLM Costs and Latency via Semantic Embedding Caching

Sajal Regmi, Chetan Phakami Pun

Large Language Models (LLMs), such as GPT, have revolutionized artificial intelligence by enabling nuanced understanding and generation of human-like text across a wide range of ap…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.