1 citations · 1 across the 14 of their papers we have counts for
12 papers · 1 filter
Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong +3
In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speake…
Typhoon ASR Real-time: FastConformer-Transducer for Thai Automatic Speech Recognition
Warit Sirichotedumrong, Adisai Na-Thalang, Potsawee Manakul +3
Large encoder-decoder models like Whisper achieve strong offline transcription but remain impractical for streaming applications due to high latency. However, due to the accessibil…
Extending Audio Context for Long-Form Understanding in Large Audio-Language Models
Yuatyong Chaichana, Pittawat Taveekitworachai, Warit Sirichotedumrong +2
Large Audio-Language Models (LALMs) are often constrained by short audio context windows, even when their text backbones support long contexts, limiting long-form audio understandi…
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
Potsawee Manakul, Woody Haosheng Gan, Michael J. Ryan +5
Current speech evaluation suffers from two critical limitations: the need and difficulty of designing specialized systems targeting individual audio characteristics, and poor corre…
FinCoT: Grounding Chain-of-Thought in Expert Financial Reasoning
Natapong Nitarach, Warit Sirichotedumrong, Panop Pitchayarthorn +3
This paper presents FinCoT, a structured chain-of-thought (CoT) prompting framework that embeds domain-specific expert financial reasoning blueprints to guide large language models…
Prior Prompt Engineering for Reinforcement Fine-Tuning
Pittawat Taveekitworachai, Potsawee Manakul, Sarana Nutanong +1
This paper investigates prior prompt engineering (pPE) in the context of reinforcement fine-tuning (RFT), where language models (LMs) are incentivized to exhibit behaviors that max…