7 papers
Multimodal Rapport Estimation in Real-World HRI
Akihiro Sakuramoto, Takato Hayashi, Ryo Miyoshi +2
Evaluating interaction quality in real-world HRI is an important challenge. If interaction quality can be estimated reliably, the results can be used to improve dialogue strategies…
Speech Signals Complement LLMs for Predicting Interpersonal Attraction in Speed Dating
Yuriko Kikuchi, Takato Hayashi, Ryusei Kimura +3
Large language models (LLMs) can predict interpersonal attraction from conversation transcripts, but it remains unclear what a speech predictor can add beyond transcript-only LLM p…
XPASS-Vis: A Dataset for Cross-Domain Personalized Image Aesthetic Assessment
Takato Hayashi, Hiroaki Takahara, Candy Olivia Mawalim +4
Personalized image aesthetic assessment (PIAA) seeks to model, at the individual level, the subjective nature of aesthetic judgments toward artworks and photographs. Aesthetic pref…
Technical Report of Nomi Team in the Environmental Sound Deepfake Detection Challenge 2026
Candy Olivia Mawalim, Haotian Zhang, Shogo Okada
This paper presents our work for the ICASSP 2026 Environmental Sound Deepfake Detection (ESDD) Challenge. The challenge is based on the large-scale EnvSDD dataset that consists of…
BERP: A Blind Estimator of Room Parameters for Single-Channel Noisy Speech Signals
Lijun Wang, Yixian Lu, Ziyan Gao +4
Room acoustical parameters (RAPs), room geometrical parameters (RGPs) and instantaneous occupancy level are essential metrics for parameterizing the room acoustical characteristics…
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan +48
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate th…