7 citations · 10 across the 9 of their papers we have counts for
13 papers
ConfusionBench: An Expert-Validated Benchmark for Confusion Recognition and Localization in Educational Videos
Lu Dong, Xiao Wang, Mark Frank +3
Recognizing and localizing student confusion from video is an important yet challenging problem in educational AI. Existing confusion datasets suffer from noisy labels, coarse temp…
InterventionLens: A Multi-Agent Framework for Detecting ASD Intervention Strategies in Parent-Child Shared Reading
Xiao Wang, Lu Dong, Ifeoma Nwogu +2
Home-based interventions like parent-child shared reading provide a cost-effective approach for supporting children with autism spectrum disorder (ASD). However, analyzing caregive…
MistyPilot: An Agentic Fast-Slow Thinking LLM Framework for Misty Social Robots
Xiao Wang, Lu Dong, Jingchen Sun +3
With the availability of open APIs in social robots, it has become easier to customize general-purpose tools to meet users' needs. However, interpreting high-level user instruction…
MCAD: Multimodal Context-Aware Audio Description Generation For Soccer
Lipisha Chaudhary, Trisha Mittal, Subhadra Gopalakrishnan +2
Audio Descriptions (AD) are essential for making visual content accessible to individuals with visual impairments. Recent works have shown a promising step towards automating AD, b…
AutoMisty: A Multi-Agent LLM Framework for Automated Code Generation in the Misty Social Robot
Xiao Wang, Lu Dong, Sahana Rangasrinivasan +3
The social robot's open API allows users to customize open-domain interactions. However, it remains inaccessible to those without programming experience. In this work, we introduce…
Towards the Synthesis of Non-speech Vocalizations
Enjamamul Hoq, Ifeoma Nwogu
In this report, we focus on the unconditional generation of infant cry sounds using the DiffWave framework, which has shown great promise in generating high-quality audio from nois…