4 papers
dOPSD: On-Policy Self-Distillation for Diffusion Language Models
Phuong Tuan Dat, Qi Li, Xinchao Wang
Diffusion large language models (dLLMs) generate text by iteratively denoising a masked sequence, offering a parallel alternative to autoregressive models, but eliciting strong rea…
VieSpeaker: A Large-Scale Vietnamese Speaker Recognition Dataset Beyond Visual Dependency
Viet Hoang Pham, Tran Trung Nguyen, Bao Thu Ho +2
Speaker recognition has advanced rapidly with large-scale training datasets, yet Vietnamese remains under-resourced, with existing corpora limited in scale and acoustic diversity.…
TACT-ful: Multi-Channel Terrain Affordance and Compliance Training for Payload-Robust Perceptive Humanoid Locomotion
Thanh Ly, Truong-Duy Dang, Chien Le +5
Foothold selection on structured terrain requires explicit reasoning about contact planarity, surface steepness, and kinematic reachability, properties not captured by a single hei…
VoxVietnam: a Large-Scale Multi-Genre Dataset for Vietnamese Speaker Recognition
Hoang Long Vu, Phuong Tuan Dat, Pham Thao Nhi +2
Recent research in speaker recognition aims to address vulnerabilities due to variations between enrolment and test utterances, particularly in the multi-genre phenomenon where the…