1 citations · 1 across the 4 of their papers we have counts for
5 papers
Typhoon ASR Real-time: FastConformer-Transducer for Thai Automatic Speech Recognition
Warit Sirichotedumrong, Adisai Na-Thalang, Potsawee Manakul +3
Large encoder-decoder models like Whisper achieve strong offline transcription but remain impractical for streaming applications due to high latency. However, due to the accessibil…
ThaiOCRBench: A Task-Diverse Benchmark for Vision-Language Understanding in Thai
Surapon Nonesung, Teetouch Jaknamon, Sirinya Chaiophat +5
We present ThaiOCRBench, the first comprehensive benchmark for evaluating vision-language models (VLMs) on Thai text-rich visual understanding tasks. Despite recent progress in mul…
Developing an Open Conversational Speech Corpus for the Isan Language
Adisai Na-Thalang, Chanakan Wittayasakpan, Kritsadha Phatcharoen +1
This paper introduces the development of the first open conversational speech dataset for the Isan language, the most widely spoken regional dialect in Thailand. Unlike existing sp…
Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast Asia
Samuel Cahyawijaya, Holy Lovenia, Joel Ruben Antony Moniz +89
Southeast Asia (SEA) is a region of extraordinary linguistic and cultural diversity, yet it remains significantly underrepresented in vision-language (VL) research. This often resu…
Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models
Kunat Pipatanakul, Potsawee Manakul, Natapong Nitarach +9
This paper introduces Typhoon 2, a series of text and multimodal large language models optimized for the Thai language. The series includes models for text, vision, and audio. Typh…