4 papers
RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark
Hongyu Jin, Siyi Wang, Yang Xiao +10
Humans process rich auditory environments through tightly integrated cognitive capabilities such as audio perception, audio reasoning, and memory. Despite recent progress in large…
RiskProp: Collision-Anchored Self-Supervised Risk Propagation for Early Accident Anticipation
Yiyang Zou, Tianhao Zhao, Peilun Xiao +9
Accident anticipation aims to predict impending collisions from dashcam videos and trigger early alerts. Existing methods rely on binary supervision with manually annotated "anomal…
Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
Wenda Zhang, Hongyu Jin, Siyi Wang +2
Speech Emotion Recognition models typically use single categorical labels, overlooking the inherent ambiguity of human emotions. Ambiguous Emotion Recognition addresses this by rep…
Multi-agent Self-triage System with Medical Flowcharts
Yujia Liu, Sophia Yu, Hongyue Jin +8
Online health resources and large language models (LLMs) are increasingly used as a first point of contact for medical decision-making, yet their reliability in healthcare remains…