3 papers
eess.AS2026
When Vocal Tone and Literal Meaning Diverge: An Acoustic-Semantic Incongruity Study for Large Audio-Language Models
Yu-Wen Chen, William Ho, Maxim Topaz +2
Affective cues across modalities may be incongruous (e.g., sarcasm or mocking praise), potentially leading to misinterpretation when relying on a single modality. Large Audio-Langu…
eess.AS2025
Hearing Health in Home Healthcare: Leveraging LLMs for Illness Scoring and ALMs for Vocal Biomarker Extraction
Yu-Wen Chen, William Ho, Sasha M. Vergez +8
The growing demand for home healthcare calls for tools that can support care delivery. In this study, we explore automatic health assessment from voice using real-world home care v…
cs.CV2025
Towards Suturing World Models: Learning Predictive Models for Robotic Surgical Tasks
Mehmet Kerem Turkcan, Mattia Ballo, Filippo Filicori +1
We introduce specialized diffusion-based generative models that capture the spatiotemporal dynamics of fine-grained robotic surgical sub-stitch actions through supervised learning…