5 papers
XPASS-Vis: A Dataset for Cross-Domain Personalized Image Aesthetic Assessment
Takato Hayashi, Hiroaki Takahara, Candy Olivia Mawalim +4
Personalized image aesthetic assessment (PIAA) seeks to model, at the individual level, the subjective nature of aesthetic judgments toward artworks and photographs. Aesthetic pref…
Technical Report of Nomi Team in the Environmental Sound Deepfake Detection Challenge 2026
Candy Olivia Mawalim, Haotian Zhang, Shogo Okada
This paper presents our work for the ICASSP 2026 Environmental Sound Deepfake Detection (ESDD) Challenge. The challenge is based on the large-scale EnvSDD dataset that consists of…
Modeling Multi-Level Hearing Loss for Speech Intelligibility Prediction
Xiajie Zhou, Candy Olivia Mawalim, Masashi Unoki
The diverse perceptual consequences of hearing loss severely impede speech communication, but standard clinical audiometry, which is focused on threshold-based frequency sensitivit…
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan +48
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate th…
Detecting Spoof Voices in Asian Non-Native Speech: An Indonesian and Thai Case Study
Aulia Adila, Candy Olivia Mawalim, Masashi Unoki
This study focuses on building effective spoofing countermeasures (CMs) for non-native speech, specifically targeting Indonesian and Thai speakers. We constructed a dataset compris…