3 papers
cs.SD2026
Grammar-Guided Hierarchical Parsing for Long-form Audio Activity Recognition
Peng Zhang, Qingyu Luo, Philip J. B. Jackson +1
Long-form audio exhibits an inherent hierarchy: fine-grained events form sub-activities, which in turn constitute higher-level activities. Prior work often models these levels sepa…
cs.SD2026
Hierarchical Activity Recognition and Captioning from Long-Form Audio
Peng Zhang, Qingyu Luo, Philip J. B. Jackson +1
Complex activities in real-world audio unfold over extended durations and exhibit hierarchical structure, yet most prior work focuses on short clips and isolated events. To bridge…
cs.CV2025
Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing
Yaru Chen, Ruohao Guo, Liting Gao +4
Weakly-supervised audio-visual video parsing (AVVP) seeks to detect audible, visible, and audio-visual events without temporal annotations. Previous work has emphasized refining gl…