4 papers · 1 filter
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
Daejin Jo, Jeeyoung Yun, Byungseok Roh +1
With the rapid progress of speech language models (SLMs), discrete speech tokens have emerged as a core interface between speech and text, enabling unified modeling across modaliti…
SGPO: Self-Generated Preference Optimization based on Self-Improver
Hyeonji Lee, Daejin Jo, Seohwan Yun +1
Large language models (LLMs), despite their extensive pretraining on diverse datasets, require effective alignment to human preferences for practical and reliable deployment. Conve…
TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
Eunseop Yoon, Hee Suk Yoon, SooHwan Eom +7
Reinforcement Learning from Human Feedback (RLHF) leverages human preference data to train language models to align more closely with human essence. These human preference data, ho…
Hexa: Self-Improving for Knowledge-Grounded Dialogue System
Daejin Jo, Daniel Wontae Nam, Gunsoo Han +4
A common practice in knowledge-grounded dialogue generation is to explicitly utilize intermediate steps (e.g., web-search, memory retrieval) with modular approaches. However, data…